JobAI is a team project that explores whether a fine-tuned language model can improve forecasts of Finnish job vacancies. It uses official Statistics Finland and KEHA data to predict vacancy counts 1, 2, and 4 quarters ahead—roughly 3, 6, and 12 months.
The forecasting pipeline downloads and checks the data, builds training examples, and compares the model with simple forecasting methods. The RAG explanation layer retrieves official bulletin passages for text answers with checked source IDs and quotes. RAG supplies context; it does not change the numeric forecasts.
The selected final project model is Qwen3-4B legacy, trained on 151,727 examples, using one shared adapter for H1/H2/H4. Its exact saved run is pinned in configs/final_model.yaml, independently of any later training run:
models/adapters/qwen-qwen3-4b__20260919T125120Z/final_adapter
The saved dataset card records 1,679 selected series, an 8-quarter history, and 151,727 training / 15,929 validation / 17,809 test examples from 12tu and 12tw. The default workflow restores this shared legacy design. All five source tables remain available, but the selected adapter was not trained on ATP or 12r5 targets. Using the saved adapter requires no retraining. Selection as the final project model is not a deployment-accuracy guarantee; the existing comparisons remain development evidence.
The figures below describe the source snapshot downloaded on 14 September 2026. Queries are defined in configs/data.yaml, and download records are saved in data/manifests/.
| Table | Coverage | Source period | Downloaded value cells | Quarterly rows |
|---|---|---|---|---|
11l1 |
National vacancies, 5 measures | 2013Q1–2026Q2 | 270 | 270 |
11n1 |
Vacancies across 5 areas, including the national total | 2013Q1–2026Q2 | 270 | 270 |
12tu |
19 provinces × 434 occupation categories | Quarter ends, 2013Q1–2026Q2 | 445,284 | 445,284 |
12tw |
19 provinces × 113 industry categories | Monthly, 2009M01–2026M07 | 453,017 | 150,290 |
12r5 |
421 geographic codes × 2 vacancy measures | Quarter ends, 2009Q1–2026Q2 | 58,940 | 58,940 |
| Total | 5 tables | 957,781 | 655,054 |
ATP survey data (11l1, 11n1) is already quarterly. For the monthly KEHA tables, the pipeline keeps March, June, September, and December observations; these are not quarterly sums. The normalized data ends at 2026Q2.
These counts include missing values: 139,438 of the quarterly rows have no numeric value. Notebook 02 preserves them and checks for duplicate keys, missing quarters, and incorrect source months. Notebook 03 then filters out unsuitable series and known duplicates. The row counts above are therefore not the number of training examples.
ATP survey estimates and KEHA registered vacancies measure different things. The pipeline keeps their source and measure labels separate.
For the selected Qwen3-4B legacy 151,727-example run, all three splits came from 12tu and 12tw, separated by time—not by table. These are occupation/province and industry/province vacancy series. Notebook 05 creates the splits; Notebook 06 uses training and validation records; Notebook 07 evaluates test records against Notebook 04's baselines.
| Split | Available examples | Used for the selected run/comparison | Purpose |
|---|---|---|---|
| Training | 151,727 | All 151,727 | Update the adapter's weights. |
| Validation | 15,929 | A 1,000-example sample during training | Monitor performance without updating weights from these answers. |
| Testing | 17,809 | Notebook 07 evaluates a matched subset; the current default is 1,000 per horizon (3,000 total) | Compare the saved adapter and baselines on the same forecast cases. |
The counts come from the saved panel dataset card; the training/validation sample settings are recorded in the earlier shared-run configuration. Earlier evaluation reports used smaller test samples, so 17,809 is the available test pool, not the number scored in every report.
The chronological boundaries are:
- Training: forecast targets must be observed by 2022Q4.
- Validation: forecast origins fall in 2023Q1–2024Q2, and their targets must also fall by 2024Q2.
- Testing: forecast origins run from 2024Q3 through the latest eligible origin, with observed targets available through 2026Q2. The last origin is 2026Q1 for H1, 2025Q4 for H2, and 2025Q2 for H4.
A forecast origin is the quarter when the prediction is made. Each example contains eight past quarters and one future target at H1, H2, or H4; overlapping history windows create many examples from each series. History may reach into an earlier split, but future answers are never included in a model's input.
ATP (11l1, 11n1) was not this adapter's validation or test dataset. Those tables remain separate benchmark/EDA series, and 12r5 remains auxiliary geographic/context data. None supplies training targets or prompt context to the selected legacy adapter. The baseline methods used for its model comparison also forecast 12tu/12tw; “baseline” means a forecasting method, not a separate source table. These roles match the earlier Notebook 04 and Notebook 05.
Run the preparation notebooks in order when rebuilding inputs; reuse saved outputs when available. Notebook 06 is only needed for new fine-tuning.
| Step | Notebook | Purpose |
|---|---|---|
| 00 | Environment check | Check packages, configuration, and API access |
| 01 | Download data | Fetch source tables and record their provenance |
| 02 | Normalize and validate | Prepare quarterly data and check its quality |
| 03 | Explore and select series | Choose suitable forecast targets |
| 04 | Evaluate baselines | Establish reference results |
| 05 | Prepare the panel | Build training, validation, and test examples |
| 06 | Fine-tune the model | Train one adapter for all three horizons |
| 07 | Evaluate the model | Compare forecasts on matching test cases |
| 08 | Collect documents | Save official publications and document metadata |
| 09 | Build the RAG index | Chunk and embed documents in persistent Chroma storage |
| 10 | Search and explain | Demonstrate saved-model forecasts and bulletin context |
| 11 | Evaluate RAG | Check retrieval quality and review citations |
| 12 | Forecast chat with RAG | Route text questions to forecasts and cited explanations |
| 13 | Streamlit dashboard | Launch interactive web interface for AI chat and figures |
The classification notebook, separate horizon profiles, and larger-model experiment are optional work outside this main sequence.
Run 12 in a GPU runtime with the full jobai/ folder, configs, saved data/metadata, pinned adapter and its evaluation manifest, and cached base model restored. Also keep data/processed/rag/chroma/ and models/embeddings/BAAI--bge-m3/. Run 08 → 09 if the index is missing; 10 and 11 are separate demo/evaluation notebooks, not prerequisites for 12. If the embedding model is absent from both saved and runtime caches, set ALLOW_EMBEDDING_DOWNLOAD = True once in 12 to save it.
Edit USER_QUESTION in the last cell and rerun it; CHAT_CONTEXT carries follow-up questions. Supported requests cover selected 12tu/12tw province/occupation or province/industry series at 1Q/2Q/4Q. The reusable backend is jobai/chat.py, with question routing in jobai/question_routing.py. Using saved assets requires no fine-tuning.
Run 13 (notebooks/13_streamlit_dashboard.ipynb) in a GPU runtime. It launches the dashboard (apps/streamlit_app.py) in the background on port 8501 and exposes it using Colab's native tunnel without needing external tools like ngrok. The dashboard includes:
- AI Agent Chatbot: Natural language forecast queries with Qwen3-4B predictions and cited official bulletins.
- Figures & Analytics: Primary national and regional vacancy charts, plus an interactive dropdown to inspect model comparison graphs.
- Theme toggle: Dark and Light modes with an animated RGB gradient border.
- Copy the full repository to Google Drive. Mount it and set the working directory and
JOBAI_REPOto its root; the notebooks default to/content/drive/MyDrive/JobAI. - To rebuild matching inputs, run 00 → 05, or 03 → 05 if normalized data is already prepared. Those notebooks share the active legacy profile. If current baseline/panel files already pass checksum and profile checks, reuse them; do not rerun them merely to select an adapter. Back up generated outputs before switching experiment profiles.
- To compare both saved adapters, skip training. Have the complete legacy 151,727 and enhanced 75,000 adapter directories and the
Qwen/Qwen3-4Bbase cache available. In a fresh GPU runtime, use 06's dependency setup if needed (only the dependency-install step), then run 07 with configs/comparison_final.yaml. Evaluation never downloads missing weights or follows the latest-training pointer.
Reproduce that experiment only by explicitly setting JOBAI_MODEL_CONFIG=configs/model.yaml before 03–06. A historical 07 comparison also needs an explicit comparison config/model profile and every listed adapter; missing archived weights are not silently skipped.
Keep existing adapters, reports, and caches. See the adapter retention guide for the required files.
Notebook 02 now has one normalization stage and publishes a checksum-bearing validation report only after all tables pass. Run the updated 02 once if notebooks 03–07 report that the old validation report lacks fingerprints; forecast notebooks 10 and 12 also accept older passed reports and check the forecast history without requiring a rerun of 02. Identical CSV contents keep the same checksums: matching 03–05 outputs remain reusable, and no adapter retraining is needed. If normalized values change, refresh 03–05. The normalization compatibility check confirmed that all five updated CSVs match the legacy-era code and the existing local CSVs byte-for-byte.
The default model is Qwen3-4B, a text-only causal language model trained with 4-bit QLoRA: a small adapter is trained while the base model stays frozen. One adapter handles all three horizons; no separate H1/H2/H4 adapters are required.
Optional new training uses all eligible mixed-horizon examples, one epoch, and a deterministic 1,000-example validation sample. Inputs contain eight quarters of history and series identifiers, without the enhanced/hierarchical prompt features. Settings live in configs/model_qwen3_4b_shared.yaml and configs/eval_qwen3_4b_legacy.yaml. The restored batch-18, checkpointing-off setup targets an A100 40GB; smaller GPUs need lower batches/checkpointing. This restores the design, not bit-for-bit historical weights or missing training metadata.
The model is compared with four baselines:
- Last value: repeat the most recent observation.
- Seasonal naive: use the corresponding quarter from the previous year.
- Ridge: predict from past quarterly values.
- Enhanced Ridge: add seasonality, trend, volatility, and zero-value features.
Training and evaluation follow time order, with no random train/test split. Training targets end at 2022Q4; validation uses origins from 2023Q1–2024Q2, with targets inside that period. Test origins start at 2024Q3 and require an observed future target, so the latest usable origin differs by horizon.
Notebook 07 defaults to legacy 151,727 versus enhanced 75,000 and four baselines on 1,000 matching cases per horizon (3,000 total). Legacy remains the pinned final model; comparison does not change that selection. Legacy uses eight past quarters, while enhanced retains the earlier comparison's declared twelve-quarter history and enhanced prompt. The original enhanced training manifest is missing, so that declaration is not independently verified training metadata. Notebook 07 reconstructs extra past quarters from checksum-verified normalized data, excludes incomplete histories for every model before sampling, and saves an exclusion report. Baselines keep their saved eight-quarter design. This compares complete saved systems, not training size alone.
Historical results remain unchanged in reports/model_evaluations/comparisons/, including comparison_c64a6d3e5c17. New results include adapter_and_baseline_metrics.csv (both adapters and baselines) and current_vs_previous_adapters.csv (legacy versus enhanced), with sMAPE, MAE, RMSE, MASE and parsing coverage. No retraining is needed; if the input checks report stale baseline/panel files, refresh 03–05 first.
The existing test cases have been used repeatedly during development. A final performance claim needs untouched future data. Historical runs also differ in training data and prompts, and some original training records are missing; the comparison reports record those limitations.
MLProject-JobAI/
├── notebooks/ # ordered pipeline and optional experiments
├── apps/ # Streamlit interactive web dashboard
├── jobai/ # shared forecasting and model helpers
├── configs/ # data, evaluation, training, and RAG settings
├── data/raw/ # saved PxWeb responses
├── data/processed/ # normalized data and forecasting examples
├── data/manifests/ # source records and dataset/run metadata
├── reports/ # metrics, predictions, and figures
├── models/ # base-model caches, adapters, and checkpoints
├── tests/ # checks for the shared forecasting helpers
├── requirements.txt # general project dependencies
└── requirements-train-colab.txt # pinned training dependencies for notebook 06
Create a focused branch and keep notebook changes reproducible. Use seed 42, preserve source and run records, and keep experiment-specific settings separate from shared defaults. Before committing notebook changes, run the affected notebooks with jupyter nbconvert --execute and keep only relevant outputs. Open a PR with the validation results for review.
The retained reports identify nine distinct trained runs, not just six experiment groups. Links identify the exact run; reevaluating the same adapter does not count as another training run.
| Model and design | Horizons | Training examples | Short outcome / saved evidence |
|---|---|---|---|
| Qwen3-4B legacy | H1/H2/H4 shared | 6,000 | Initial small-data run. |
| Qwen3-4B legacy | H1/H2/H4 shared | 151,727 | Selected final model. |
| Qwen3-4B enhanced | H1/H2/H4 shared | 6,000 | Small-data enhanced-input trial. |
| Qwen3-4B enhanced | H1/H2/H4 shared | 75,000 | Larger enhanced-input trial. |
| Qwen3-4B enhanced | H1 only | 10,000 | One-quarter specialization trial. |
| Qwen3-4B enhanced | H2 only | 10,000 | Two-quarter specialization trial. |
| Qwen3-4B enhanced | H4 only | 10,000 | Four-quarter specialization trial. |
| Qwen3-4B enhanced | H1 only | 60,000 | Larger H1 specialization trial. |
| Qwen3.5-4B shared enhanced | H1/H2/H4 shared | 75,000 | Evaluated against retained Qwen3-4B adapters; not selected. |
These are part of the development history, but no completed evaluation for them was found in the retained run reports. A configuration file alone does not prove training finished.
| Model / project branch | What was tried | Evidence |
|---|---|---|
| Qwen3.5-9B | Early preferred QLoRA model; later replaced in the active workflow. | Historical configuration, 0319d2f. |
| Qwen3-8B | Text-only alternative before the Qwen3-4B runs. | Historical configuration, 77b22bf. |
Archived 27B experiment, labelled Qwen3.8-27B in the project |
Separate version-2 notebook and large-model configuration; not the selected model. | Archived configuration. The name here records the project label, not independent verification of a released model ID. |
We also tested different batch sizes, gradient accumulation, checkpointing, validation budgets, caching and GPU setups. Those setup changes, interrupted attempts, and repeated evaluations are not counted as distinct completed training runs unless a separate run record supports them.
Runs differ in training budget, features, and sometimes evaluation cohorts. Older pilot_passed/pilot_failed labels apply only to that report's sample and criteria; they are not permanent rankings or deployment approval. Some original training manifests are missing, so evaluation records remain the available evidence. This history does not claim architecture-only superiority or under-10% sMAPE. No adapter or historical report is removed by documenting the history.