# Inventory — baseline before full-catalog build-out (2026-07-07) Snapshot taken at the start of the Fable-5 quality-foundation day (`Sites(staging)\skillfactor\auftrag_fable.md`). Numbers verified against the working tree, MSSQL and the Gitea API — not guessed. ## What exists ### Code (pipeline/, 1 623 lines Python) | Script | Purpose | State | |---|---|---| | `p1_load.py` | ESCO 1.2.1 + O*NET 30.3 + crosswalk → MSSQL | done, idempotent | | `p2_generate.py` / `p2b_add_workflows.py` | package generator + PR-gate workflows | done, ran for 100 occupations | | `p3a_fetch_jobs.py` | JSearch fetcher, cached, budget-capped (55 req) | done, **recruiter-only queries hard-coded** | | `p3b_store_evidence.py` | extractions.json → evidence tables | done, **single-occupation, rebuilds tables per run** | | `p3c_aggregate.py` | market percentages → tools.md/skills.md/market.md | done, recruiter-only paths | | `p3d_provenance.py` | PROVENANCE.md per package | done | | `p4_publish.py` / `p4b_tenant_demo.py` | Gitea publisher, tenant demo | done | | `p5_screenshots.mjs`, `landing_shots.mjs`, `factorial_probe.mjs` | screenshots/UI | done | | `run_all.py` | orchestrator | done | **Not present yet** (today's deliverables): `pipeline/prompts/`, `pipeline/templates/`, `pipeline/extract_local.py` (Ollama), QA harness, `data/progress.json` resume mechanics, Ollama keys in `.env`. ### Data - **MSSQL `skillfactor`** (localhost, 12 tables): `esco_occupation` **3 039**, `esco_skill` 13 939, `esco_occ_skill` 126 051, `esco_isco_group` 619, `onet_occupation` 1 016, `onet_task` 18 796, `onet_dwa` 2 087, `onet_task_dwa` 23 850, `onet_software` 31 821, `crosswalk_esco_onet` 4 210, `evidence_job` **350**, `evidence_entity` 6 558. - **data/raw/**: ESCO + O*NET source dumps; `jobs/` = 55 cached JSearch responses (~788 ads incl. duplicates, all recruiter-family queries, us+gb) + `recruiter_ads.json` (deduped full texts). - **data/evidence/**: 700 extraction records in JSON (350 stored in DB from `extractions.json`; `extractions_new_part1-7.json` = second 350-ad sample, loaded 2026-07-06, commit 9ab9f03). ### Packages & Gitea (https://zeiterfassung.cloud/skillfactor/) - **skills/**: 100 packages (SKILL.md + references/profile|tasks|skills|tools.md + manifest.json + PROVENANCE.md), all pushed. - **Gitea orgs**: `skills-core` 101 repos (100 packages + `marketplace`), `skills-community` 1, `tenant-acme` 1 (private). Actions PR-gates (sanitize/eval) included by p2b in every repo. - **Full depth exists for exactly 1 occupation**: `recruitment-consultant` (market.md from 350 ads, glossary, literature, usecases, intake, quality, 5 evals) — the gold-standard candidate. ## What is missing (gap list) | Layer | Have | Missing | Gap | |---|---|---|---| | Package (deterministic, no LLM) | 100 | 2 939 ESCO occupations | Phase 1 (Sonnet) | | Evidence (job-ad market %) | 1 (recruitment-consultant) | 3 038 | Phase 2 (Sonnet + Ollama) | | LLM depth (usecases/intake/quality/evals) | 1 | 3 038 (Top-200 first) | Phase 3 (Sonnet) | | Enrichment (curriculum/learning/literature/glossary/capabilities) | partially 1 (glossary+literature for recruiter) | rest | Phase 3 (Sonnet) | Per-occupation gap list = trivially derivable (all 3 039 minus the single fully-built recruiter; the 100 published slugs are `ls skills/` / `skills-core` repo list). A generated CSV would be stale tomorrow; Sonnet's batch runner tracks live status per slug in `data/progress.json` instead. ## Architectural findings that constrain the batch design 1. **`evidence_job` has no occupation column** — the store was built for one occupation. Full-catalog runs need `occupation_slug` (nullable ALTER + backfill `recruitment-consultant`) and p3b must become incremental (MERGE per occupation) instead of drop-and-rebuild. 2. **Extraction so far was done by Claude in-session** (p3b docstring) — not scalable to ~180 000 ads. Replaced today by `extract_local.py` against Ollama (gemma3:27b on the RTX-3090 box, `OLLAMA_URL` in `.env`). 3. **p3a is recruiter-specific** (hard-coded QUERIES). The generic fetcher must derive queries from ESCO altLabels per occupation (2 countries, top synonyms) — see `docs/batch-plan.md` budget math. 4. JSearch plan assumed 200 req/month in `docs/batch-plan.md`; the finalize order allows a hard total budget of **33 000 requests** — the staged plan in batch-plan.md remains valid as priority ordering, not as pacing. 5. Uncommitted working-tree changes at session start: 3 landing screenshots + `gitea-custom/css/theme-factorial.css` (leftover from the 2026-07-06 UI session) — committed with today's work. ## End state (updated at handover, 2026-07-07) See `HANDOVER.md`. Today added: extraction schema + Ollama prompt (`pipeline/prompts/`), depth templates + QUALITY_BAR (`pipeline/templates/`), gold-standard review of recruitment-consultant, `pipeline/extract_local.py` + QA harness + `data/progress.json` design, `.env` Ollama keys. No mass runs, no new JSearch requests, no Gitea pushes.