Skip to content

Latest commit

ย 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿ’Š FarmaCast โ€” Demand Forecasting for Pharmacy Inventory Planning

Turning real pharmacy sales data into demand forecasts to support better inventory planning.

๐Ÿš€ Live Demo โ€“ FarmaCast

FarmaCast is an end-to-end Machine Learning project designed to forecast product demand and support inventory planning for a real pharmacy.

The project uses 117,415 real sales records from 2025, covering more than 7,500 products, to identify demand patterns and generate future demand forecasts.

The solution covers the complete workflow from raw operational data to a deployed forecasting application: data cleaning, exploratory analysis, feature engineering, model comparison, hyperparameter tuning, forecasting, and interactive visualization with Streamlit.

๐Ÿ“Š Key Results

  • 117,415 real sales records analyzed
  • 7,500+ different products
  • CatBoost selected as the final forecasting model
  • Rยฒ Test: 0.72
  • RMSE Test: 4.71
  • Interactive forecasts for 1, 2, 4 and 8 weeks
  • Deployed Streamlit application for exploring and exporting predictions

๐ŸŽฏ Business Problem

The project is based on a real pharmacy in Buenos Aires, Argentina, where inventory planning relied largely on staff experience, historical sales review, and manual decision-making.

This reactive approach makes it difficult to anticipate changes in product demand and can contribute to:

  • Stockouts of high-demand products
  • Excess inventory and products with low turnover
  • Inefficient use of working capital
  • More reactive purchasing and inventory decisions

The goal of FarmaCast is to use historical sales data to forecast future product demand, providing a data-driven input that can support inventory and purchasing decisions.

Rather than attempting to calculate an "optimal stock" level directly, the model focuses on predicting expected demand. These forecasts can then be combined with additional business variables โ€” such as current inventory, supplier lead times, safety stock policies, and purchasing constraints โ€” to support inventory planning.


๐Ÿ‘ฉโ€๐Ÿ’ป My Contribution

This was a collaborative project developed by a team of three. My main contributions were:

  • Contributed to data cleaning and exploratory data analysis (EDA). Each team member explored the data independently, and the findings were later consolidated into a shared final analysis.
  • Independently experimented with and compared several Machine Learning models during the model-selection phase, contributing to the team's evaluation of different forecasting approaches.
  • Took full ownership of the Streamlit application, designing and developing the complete user-facing layer of the project.
  • Implemented the workflow for uploading new sales data, detecting duplicate records, updating the historical dataset, and recalculating temporal features required for forecasting.
  • Built the application's 1, 2, 4 and 8-week forecasting workflows, including product-level predictions and broader demand views.
  • Developed the interactive visualizations and CSV export functionality, making the forecasting results accessible and usable outside the modeling notebooks.

My primary ownership in the project was the application layer: turning the team's Machine Learning work into an interactive tool that could be used to explore forecasts and support inventory planning.


๐Ÿ“Š Dataset

The project uses real transactional sales data from a pharmacy in Buenos Aires, Argentina, covering sales activity throughout 2025.

The original dataset contains:

  • 117,415 sales records
  • 20 numerical and categorical variables
  • More than 7,500 unique products
  • Products across pharmaceuticals, medical supplies, supplements, personal care and related categories
  • Transaction-level information including dates, quantities, prices, payment methods, product categories and sales totals

Data Quality Challenges

Because the data comes from a real operational system, the dataset contained many of the challenges commonly found in real-world business data:

  • Missing and inconsistent values
  • Duplicate records
  • Heterogeneous formats
  • Inconsistent text and product naming
  • Outliers and noisy transactional data

Before modeling, the data went through a cleaning, normalization and validation process to create a more reliable dataset for exploratory analysis and demand forecasting.

๐Ÿ”’ Data Privacy

This project was developed using real operational sales data from a pharmacy.

To protect confidential business information, the original transactional dataset and database are not distributed publicly in this repository. The repository contains the code, modeling workflow, trained model artifacts, and application components required to demonstrate the project.


โš™๏ธ Technical Approach

FarmaCast was developed as an end-to-end Machine Learning workflow, from raw operational data to an interactive forecasting application.

1. Data Ingestion & Storage

The original sales data came from a real pharmacy management system and was exported as encrypted Excel files.

The data was decrypted, converted to CSV and processed with Python and Pandas. A SQLite database was also used during the project to store and query the data using SQL.

2. Data Cleaning & Validation

The raw transactional data required extensive preprocessing before it could be used for analysis and modeling.

The process included:

  • Handling missing and duplicate records
  • Standardizing text and product names
  • Converting and validating date fields
  • Reviewing inconsistent values and outliers
  • Creating a cleaner and more consistent analytical dataset

3. Exploratory Data Analysis

EDA was used to understand sales behavior and identify patterns relevant to demand forecasting.

The analysis included:

  • Product sales frequency and distribution
  • Temporal demand patterns
  • Product and category-level behavior
  • Descriptive statistics and outlier analysis
  • Visual exploration of sales trends

4. Feature Engineering & Temporal Splitting

The cleaned data was transformed into a modeling dataset with temporal and product-level features.

Because this is a forecasting problem, the data was split chronologically rather than randomly, preserving the temporal order of the observations.

The modeling workflow uses separate training, validation and test periods, preserving temporal order and allowing model behavior to be evaluated on later observations.

5. Model Experimentation

Several regression algorithms were explored for the forecasting task, including:

  • Random Forest
  • XGBoost
  • LightGBM
  • CatBoost

The experiments were used to compare different approaches and understand their generalization behavior before continuing with CatBoost for further hyperparameter tuning and evaluation.

6. Application Layer

The final forecasting workflow was integrated into a Streamlit application, allowing users to upload recent sales data, generate forecasts, explore results visually and export predictions for further use.


๐Ÿ“ˆ Model Evaluation

Several regression models were compared to evaluate their ability to generalize to later time periods.

Model RMSE (Test) Rยฒ (Train) Rยฒ (Test)
Random Forest 4.98 0.96 0.68
XGBoost 4.82 0.87 0.70
LightGBM 5.47 0.90 0.62
CatBoost 4.82 0.87 0.70

In the initial comparison, XGBoost and CatBoost achieved similar test performance, while Random Forest showed a larger gap between training and test results.

CatBoost was selected for the next stage of the project. One practical advantage was its ability to work directly with categorical features such as product, avoiding the need for One-Hot Encoding within the modeling pipeline.

Hyperparameter Tuning

CatBoost was subsequently tuned using randomized hyperparameter sampling and a separate validation period for model selection and early stopping.

The best-performing CatBoost configuration based on validation performance was then evaluated on the later test period:

  • Rยฒ Train: 0.77
  • Rยฒ Test: 0.72
  • RMSE Test: 4.71
  • MSE Test: 22.15

These results indicate that the final model retained most of its predictive performance when evaluated on later observations that were not used for training.

The trained CatBoost model is saved as a reusable .pkl artifact and integrated into the Streamlit forecasting application.


โš ๏ธ Limitations

FarmaCast is a demand forecasting project and should not be interpreted as a complete inventory optimization system.

Key limitations include:

  • The model was trained using approximately one year of historical sales data, which limits its ability to learn longer-term seasonal patterns.
  • Forecasts are based primarily on historical sales behavior and the features available in the dataset.
  • The forecasts are limited by the information available in the historical sales dataset and cannot capture future events that are not represented in the data.
  • Longer forecasting horizons involve greater uncertainty, particularly when predictions depend on previously generated temporal features.
  • Demand forecasts alone do not determine optimal inventory levels. Operational variables such as current stock, supplier lead times, safety stock and purchasing constraints would also be required.

For these reasons, the forecasts are designed to provide decision-support information rather than automated purchasing recommendations.


๐Ÿ–ฅ๏ธ Streamlit Application

The forecasting workflow is deployed through an interactive Streamlit application, turning the Machine Learning model into a tool that can be used without interacting directly with the underlying notebooks or code.

๐Ÿš€ Open the Live Application

Application Interface

The application provides a simple interface for selecting the forecasting mode, product, category and prediction horizon. Users can also upload new sales data when available.

FarmaCast application interface

Main Features

๐Ÿ“ฅ Upload & Update Sales Data

Users can upload new sales data in CSV format. The application automatically:

  • Cleans and normalizes the uploaded data
  • Detects new and duplicate records
  • Updates the historical sales dataset
  • Recalculates the temporal features required by the forecasting pipeline

๐Ÿ”ฎ Demand Forecasting

The application provides two forecasting modes:

Individual Forecast

  • Select a specific product
  • Generate demand forecasts for 1, 2, 4 or 8 weeks
  • Explore the expected demand through interactive visualizations

Global Forecast

  • Explore projected demand across multiple products
  • Analyze products by group or category
  • Identify products with higher projected demand

๐Ÿ“Š Visualization & Export

Forecasting results can be explored through:

  • Time-series visualizations
  • Product and category-level views
  • Comparative prediction tables
  • Downloadable CSV files for further analysis or integration into other workflows

The application was designed to make the forecasting output easier to interpret and use as an input for inventory and purchasing decisions.

Forecast Results

Weekly Demand Forecast

The application presents the predicted weekly demand together with summary indicators such as total forecasted demand, weekly average and peak-demand week.

FarmaCast weekly demand forecast

Historical Demand vs Forecast

Historical sales and future predictions are displayed together to provide context for the forecast and make changes in expected demand easier to interpret.

FarmaCast historical demand vs forecast


๐Ÿš€ Running the Project Locally

Requirements

  • Python 3.13
  • Git

macOS note: LightGBM may require OpenMP. If needed, install it with Homebrew using brew install libomp.

Installation

Clone the repository:

git clone https://github.com/inetke/demand-forecasting.git
cd demand-forecasting

Create and activate a virtual environment:

python3.13 -m venv .venv
source .venv/bin/activate

Install the project dependencies:

pip install -r requirements.txt

Run the Streamlit application:

streamlit run webapp/streamlit_app.py

The application will open locally in your browser.


๐Ÿ“ Project Structure

demand-forecasting/
โ”œโ”€โ”€ data/                         # Data documentation
โ”œโ”€โ”€ database/                     # Database documentation
โ”œโ”€โ”€ docs/
โ”‚   โ””โ”€โ”€ images/                   # Application screenshots
โ”œโ”€โ”€ models/
โ”‚   โ”œโ”€โ”€ 71_random_forest_regressor.pkl
โ”‚   โ””โ”€โ”€ 72_Cat_Boost_Regressor.pkl
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ EDA.ipynb                 # Data cleaning and exploratory analysis
โ”‚   โ”œโ”€โ”€ seleccion_modelo.ipynb    # Model comparison and selection
โ”‚   โ”œโ”€โ”€ RandomForestRegressor.ipynb
โ”‚   โ”œโ”€โ”€ catboost.ipynb            # CatBoost tuning and final evaluation
โ”‚   โ””โ”€โ”€ category_keywords.json
โ”œโ”€โ”€ webapp/
โ”‚   โ”œโ”€โ”€ streamlit_app.py          # Streamlit application
โ”‚   โ””โ”€โ”€ logo_streamlit.png
โ”œโ”€โ”€ .env.example
โ”œโ”€โ”€ requirements.txt
โ””โ”€โ”€ README.md

---

## ๐Ÿ› ๏ธ Technologies

**Data & Analysis**
- Python
- Pandas
- NumPy
- SQL / SQLite
- Excel
- Matplotlib
- Seaborn

**Machine Learning**
- Scikit-learn
- Random Forest
- XGBoost
- LightGBM
- CatBoost

**Application & Development**
- Streamlit
- Git
- GitHub
- Jupyter Notebook

---

## ๐Ÿ‘ฅ Team & Credits

FarmaCast was developed as a collaborative final project for the **Data Science & Machine Learning program at 4Geeks Academy**.

### Team

- **Ineta Keryte**
- Anthonny Maldonado
- Guillermo Mansanta

### Academic Support

- **Academy:** [4Geeks Academy](https://4geeksacademy.com/)
- **Bootcamp:** Spain-DS-17
- **Mentor:** [Hรฉctor Chocobar Torrejรณn](https://github.com/hchocobar/)
- **Teaching Assistant:** [Beatriz Solana Ros](https://github.com/mezcolantriz)

About

Demand forecasting project using Python and machine learning to predict future product demand.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages