Files
skillfactor-pipeline/docs/evidence-report.md
skillfactor-pipeline af05383bb1 chore: v1 baseline - full 3,039-package catalog with AI-skill enrichment and O*NET backfill
Snapshot before the quality program (relevance gates, tiered mapping,
QA linter). v1 is the immutable before/after reference; evidence crawl
was at ~175/3039 occupations when tagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDKeXvpT6tENSvyQGLV1Uq
2026-07-09 05:32:49 +02:00

67 lines
1.5 KiB
Markdown

# Job-ad evidence — coverage report
_Generated 2026-07-08 by `pipeline/evidence_report.py`. Re-run any time for the current state._
## Coverage
| Evidence status | Occupations |
|---|---|
| pending | 3039 |
| **total catalog** | **3039** |
- Occupations with ads in the evidence store: **1**
- Thin markets (fewer than 5 usable ads): **0**
- Failed for other reasons: **0**
## Volume
| Metric | Value |
|---|---|
| Job ads stored | 350 |
| Distinct employers | 309 |
| Extracted entities | 6,558 |
| — responsibilities | 2,502 |
| — hard_skills | 2,244 |
| — methods | 1,009 |
| — tools | 803 |
## API budget
- JSearch requests spent: **0 / 33,000** (0.0 % of the lifetime budget)
- Every request is counted BEFORE the HTTP call (`progress.spend_request`, hard cap).
## Distributions
**Seniority** (per ad):
| Seniority | Ads |
|---|---|
| mid | 133 |
| junior | 102 |
| senior | 57 |
| n/a | 31 |
| lead | 27 |
**Country**:
| Country | Ads |
|---|---|
| us | 308 |
| gb | 42 |
**Top 20 occupations by ads:**
| Occupation | Ads |
|---|---|
| recruitment-consultant | 350 |
## Extraction engines
| Engine | Ads extracted | Share | Role |
|---|---|---|---|
| Claude (in-session) | 350 | 100.0 % | flagship package (recruitment-consultant), prompt/validator design, 2 % QA spot checks on every batch |
| Ollama gemma3:27b (RTX-3090 box) | 0 | 0.0 % | mass extraction of the full catalog (`extract_local.py`, strict JSON schema, validator with final word) |
_Thin-market occupations (kept for transparency):_