feat: big-data-archive-librarian skill package v0.1.0

This commit is contained in:
skillfactor-pipeline
2026-08-14 18:56:44 +02:00
commit fbb76bb34c
9 changed files with 862 additions and 0 deletions

31
PROVENANCE.md Normal file
View File

@@ -0,0 +1,31 @@
# Data provenance — big-data-archive-librarian
Where the content of this skill package comes from, counted by
content items (tasks, competences, tools, evidence entries, curated
knowledge). Rendered live by Gitea:
```mermaid
%%{init: {'theme':'base','themeVariables':{'pie1':'#f9a825','pie2':'#1e88e5','pie3':'#ff355e','pie4':'#d97757','pie5':'#8e24aa','pieOuterStrokeWidth':'0px','pieSectionTextColor':'#fff'}}}%%
pie showData
title Content sources — big-data-archive-librarian
"ESCO (occupation & competences)" : 74
"O*NET (tasks & tools)" : 89
"Job boards (market evidence)" : 34
"Anthropic official Claude skills" : 5
"External AI skill packs (mapped)" : 141
```
| Source | Items | Share | Files |
|---|---|---|---|
| ESCO (occupation & competences) | 74 | 21.6 % | references/profile.md, references/skills.md |
| O*NET (tasks & tools) | 89 | 25.9 % | references/tasks.md, references/tools.md |
| Job boards (market evidence) | 34 | 9.9 % | references/market.md (full report) + "Market evidence" headline sections |
| Wikipedia & AI expert curation | 0 | 0.0 % | glossary, literature, usecases, intake, quality, evals/ |
| Anthropic official Claude skills | 5 | 1.5 % | references/ai-skills.md, section "anthropics/skills" (official Claude Code skills) |
| External AI skill packs (mapped) | 141 | 41.1 % | references/ai-skills.md (per-source attribution inside) |
| Stack Exchange practitioner Q&A (CC-BY-SA) | 0 | 0.0 % | references/practitioner-qa.md (per-entry attribution inside) |
Licensing: O*NET (USDOL/ETA, CC BY 4.0) · ESCO (© European Union) ·
job-ad evidence via official APIs (JSearch/Adzuna) · Wikipedia content
paraphrased with source URLs — never copied · external AI skills are
linked, not copied (Apache-2.0/MIT/source-available, see ai-skills.md).

61
SKILL.md Normal file
View File

@@ -0,0 +1,61 @@
---
name: big-data-archive-librarian
description: "Occupational skill for the role 'big data archive librarian' (also: digital documentation archivist, archive librarian, documentation archivist, computer tape librarian, digital archivists). Use when the user asks for typical big data archive librarian work such as: Open and close library during specified hours and secure library equipment, such as computers and audio-visual equipment.; Perform clerical activities, such as answering phones, sorting mail, filing, typing, word processing, and photocopying and mailing out material.; Schedule, supervise, and train clerical workers, volunteers, student assistants, and other library employees."
---
# Big Data Archive Librarian
Big data archive librarians classify, catalogue and maintain libraries of digital media. They also evaluate and comply with metadata standards for digital content and update obsolete data and legacy systems.
## Core workflow
1. Open and close library during specified hours and secure library equipment, such as computers and audio-visual equipment.
2. Perform clerical activities, such as answering phones, sorting mail, filing, typing, word processing, and photocopying and mailing out material.
3. Schedule, supervise, and train clerical workers, volunteers, student assistants, and other library employees.
4. Maintain library equipment, such as photocopiers, scanners, and computers, and instruct patrons in proper use of such equipment.
5. Manage reserve materials by placing items on reserve for library patrons, checking items in and out of library, and removing out-of-date items.
6. Lend, reserve, and collect books, periodicals, videotapes, and other materials at circulation desks and process materials for inter-library loans.
7. Repair books using mending tape, paste, and brushes or prepare books to be sent to a bindery for repair.
8. Prepare library statistics reports.
## How to use this skill
- Read [references/profile.md](references/profile.md) for the occupation profile and scope.
- Consult [references/tasks.md](references/tasks.md) for the full task and activity inventory.
- Check [references/skills.md](references/skills.md) for essential vs. optional competences.
- Check [references/tools.md](references/tools.md) for the software commonly used in this role.
- See [references/ai-skills.md](references/ai-skills.md) — matched external AI agent skills (per-source attribution).
## Key competences (essential)
- analyse big data
- business intelligence
- comply with legal regulations
- data extraction, transformation and loading tools
- data models
- database
- database development tools
- database management systems
- digital curation
- digital data processing
- digitization
- maintain data entry requirements
- maintain database performance
- maintain database security
- manage archive users guidelines
## Hot technologies
- Microsoft Access
- Adobe Acrobat
- Microsoft Outlook
- Adobe Photoshop
- C++
- Microsoft Office software
- Microsoft Windows
- Microsoft PowerPoint
- Microsoft Excel
- Microsoft Word
---
*Sources: ESCO v1.2.1 (http://data.europa.eu/esco/occupation/4e4a1aa1-ca53-406e-9ca6-889d4621a0ed), O*NET 30.3 (43-4121.00). See manifest.json for licensing/attribution.*

202
manifest.json Normal file
View File

@@ -0,0 +1,202 @@
{
"name": "big-data-archive-librarian",
"title": "big data archive librarian",
"version": "0.1.0",
"layer": "core",
"language": "en",
"generated": "2026-07-07",
"ids": {
"esco_uri": "http://data.europa.eu/esco/occupation/4e4a1aa1-ca53-406e-9ca6-889d4621a0ed",
"esco_code": "3433.2",
"isco_group": "3433",
"onet_soc": "43-4121.00",
"crosswalk_match": "closeMatch"
},
"sources": [
{
"name": "ESCO",
"version": "1.2.1",
"url": "https://esco.ec.europa.eu/"
},
{
"name": "O*NET",
"version": "30.3",
"url": "https://www.onetcenter.org/",
"license": "CC BY 4.0"
}
],
"attribution": "This package includes information from the O*NET Database (v30.3) by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA), CC BY 4.0. skillfactor is not endorsed by USDOL/ETA. ESCO data (v1.2.1) (c) European Union, used per the ESCO download conditions: https://esco.ec.europa.eu/en/use-esco/download",
"counts": {
"tasks": 32,
"dwas": 36,
"skills_essential": 23,
"skills_optional": 50,
"software": 20
},
"enrichment_ai_skills": {
"generated": "2026-07-14",
"method": "deterministic mapping (ISCO prefix + title/competence keywords)",
"sources": {
"anthropics/skills": {
"repo": "https://github.com/anthropics/skills",
"commit": "f6656c1",
"license": "Apache-2.0; the document skills (docx/pdf/pptx/xlsx) are source-available \u2014 see the LICENSE.txt in the upstream skill folder",
"skills": 2
},
"wshobson/agents": {
"repo": "https://github.com/wshobson/agents",
"commit": "6fd3247",
"license": "MIT (c) Seth Hobson",
"skills": 12
},
"NVIDIA/skills": {
"repo": "https://github.com/NVIDIA/skills",
"commit": "153b14b",
"license": "CC-BY-4.0 (skills/docs), Apache-2.0 (code)",
"skills": 12
},
"veniceai/skills": {
"repo": "https://github.com/veniceai/skills",
"commit": "de089fa",
"license": "MIT",
"skills": 6
},
"ConardLi/garden-skills": {
"repo": "https://github.com/ConardLi/garden-skills",
"commit": "fbd6453",
"license": "MIT",
"skills": 1
},
"a5c-ai/babysitter": {
"repo": "https://github.com/a5c-ai/babysitter",
"commit": "44a5d58b",
"license": "MIT",
"skills": 8
},
"nowork-studio/NotFair": {
"repo": "https://github.com/nowork-studio/NotFair",
"commit": "99bda4b",
"license": "MIT",
"skills": 1
},
"K-Dense-AI/claude-scientific-skills": {
"repo": "https://github.com/K-Dense-AI/claude-scientific-skills",
"commit": "4d97e29",
"license": "MIT",
"skills": 1
},
"foryourhealth111-pixel/Vibe-Skills": {
"repo": "https://github.com/foryourhealth111-pixel/Vibe-Skills",
"commit": "34429a8",
"license": "Apache-2.0",
"skills": 4
},
"K-Dense-AI/scientific-agent-skills": {
"repo": "https://github.com/K-Dense-AI/scientific-agent-skills",
"commit": "4d97e29",
"license": "MIT",
"skills": 1
},
"davila7/claude-code-templates": {
"repo": "https://github.com/davila7/claude-code-templates",
"commit": "fa79251",
"license": "MIT",
"skills": 4
},
"jawwadfirdousi/agent-skills": {
"repo": "https://github.com/jawwadfirdousi/agent-skills",
"commit": "962ffed",
"license": "no explicit license \u2014 referenced by link only",
"skills": 1
},
"davepoon/buildwithclaude": {
"repo": "https://github.com/davepoon/buildwithclaude",
"commit": "3c94e0c",
"license": "MIT",
"skills": 6
},
"jeremylongshore/claude-code-plugins-plus-skills": {
"repo": "https://github.com/jeremylongshore/claude-code-plugins-plus-skills",
"commit": "e112938a",
"license": "MIT",
"skills": 6
},
"secondsky/claude-skills": {
"repo": "https://github.com/secondsky/claude-skills",
"commit": "c4889f6",
"license": "MIT",
"skills": 1
},
"ruvnet/claude-code-flow": {
"repo": "https://github.com/ruvnet/claude-code-flow",
"commit": "73914bd",
"license": "MIT",
"skills": 1
},
"ruvnet/ruflo": {
"repo": "https://github.com/ruvnet/ruflo",
"commit": "73914bd",
"license": "MIT",
"skills": 1
},
"giuseppe-trisciuoglio/developer-kit": {
"repo": "https://github.com/giuseppe-trisciuoglio/developer-kit",
"commit": "306f428",
"license": "MIT",
"skills": 2
},
"brycewang-stanford/Auto-Empirical-Research-Skills": {
"repo": "https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills",
"commit": "85bf545",
"license": "CC-BY-4.0",
"skills": 7
},
"wondelai/skills": {
"repo": "https://github.com/wondelai/skills",
"commit": "326b380",
"license": "MIT",
"skills": 1
},
"sgcarstrends/backend": {
"repo": "https://github.com/sgcarstrends/backend",
"commit": "7231cbc",
"license": "MIT",
"skills": 1
},
"samber/cc-skills-golang": {
"repo": "https://github.com/samber/cc-skills-golang",
"commit": "4881c01",
"license": "MIT",
"skills": 1
}
},
"total_skills": 80,
"tiers": {
"core": 66,
"adjacent": 14
}
},
"provenance": {
"items": {
"esco": 74,
"onet": 89,
"jobads": 34,
"wiki_ai": 0,
"anthropic": 5,
"ai_skills": 141,
"stackx": 0
},
"share_percent": {
"esco": 21.6,
"onet": 25.9,
"jobads": 9.9,
"wiki_ai": 0.0,
"anthropic": 1.5,
"ai_skills": 41.1,
"stackx": 0.0
},
"method": "content items per source category"
},
"collar": "white",
"computer_work": true
}

274
references/ai-skills.md Normal file
View File

@@ -0,0 +1,274 @@
# External AI agent skills — big-data-archive-librarian
Proven, publicly available AI agent skills mapped to this occupation.
Nothing is copied from the sources: every entry is a name, a one-line
summary and a link to the upstream skill package. Each section names
its source repository, commit, license and retrieval date.
**Tiers:** `core` = the skill directly exercises a top market hard
skill, tool or method (from gated job-ad evidence) or an essential
ESCO competence of this occupation; `adjacent` =
plausibly useful, secondary. Entries are capped at 12 per source
and 80 in total per occupation (core first,
strongest matches survive); everything beyond the caps is excluded
and logged in the pipeline audit trail, not in this package.
_Matched deterministically (ISCO group + title/competence keywords,
tiered against market evidence + ESCO essentials) by
`pipeline/p5_enrich_ai_skills.py` on 2026-07-14._
## Source: anthropics/skills
- Repository: [https://github.com/anthropics/skills](https://github.com/anthropics/skills) (commit `f6656c1`, retrieved 2026-07-14)
- License: Apache-2.0; the document skills (docx/pdf/pptx/xlsx) are source-available — see the LICENSE.txt in the upstream skill folder
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `docx` | adjacent | Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to … | [source](https://github.com/anthropics/skills/tree/main/skills/docx) |
| `pdf` | adjacent | Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating … | [source](https://github.com/anthropics/skills/tree/main/skills/pdf) |
## Source: wshobson/agents
- Repository: [https://github.com/wshobson/agents](https://github.com/wshobson/agents) (commit `6fd3247`, retrieved 2026-07-14)
- License: MIT (c) Seth Hobson
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `dbt-transformation-patterns` | core | Master dbt (data build tool) for analytics engineering with model organization, testing, documentation, and incremental strategies. Use when building data transformations, creating data models, or implementing analytics engineering best … | [source](https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/dbt-transformation-patterns) |
| `data-validation-suite-backend-security-coder (agent)` | core | Expert in secure backend coding practices specializing in input validation, authentication, and API security. Use PROACTIVELY for backend security implementations or security code reviews. | [source](https://github.com/wshobson/agents/tree/main/plugins/data-validation-suite/agents/backend-security-coder.md) |
| `ml-pipeline-workflow` | core | Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment. Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows. | [source](https://github.com/wshobson/agents/tree/main/plugins/machine-learning-ops/skills/ml-pipeline-workflow) |
| `airflow-dag-patterns` | core | Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs. | [source](https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/airflow-dag-patterns) |
| `data-quality-frameworks` | core | Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts. | [source](https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/data-quality-frameworks) |
| `embedding-strategies` | core | Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains. | [source](https://github.com/wshobson/agents/tree/main/plugins/llm-application-dev/skills/embedding-strategies) |
| `spark-optimization` | adjacent | Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines. | [source](https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization) |
| `recsys-pipeline-architect` | adjacent | Design composable recommendation, ranking, and feed pipelines using the six-stage Source→Hydrator→Filter→Scorer→Selector→SideEffect framework popularized by xAI's open-sourced X For You algorithm. Use when building any system that picks … | [source](https://github.com/wshobson/agents/tree/main/plugins/machine-learning-ops/skills/recsys-pipeline-architect) |
| `llm-evaluation` | adjacent | Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks. | [source](https://github.com/wshobson/agents/tree/main/plugins/llm-application-dev/skills/llm-evaluation) |
| `prompt-engineering-patterns` | adjacent | This skill should be used when the user asks to "optimize a prompt", "improve prompt performance", "design a prompt template", "write better prompts", "debug prompt issues", "use chain-of-thought", "structured prompting", "few-shot … | [source](https://github.com/wshobson/agents/tree/main/plugins/llm-application-dev/skills/prompt-engineering-patterns) |
| `hybrid-search-implementation` | adjacent | Combine vector and keyword search for improved retrieval. Use when implementing RAG systems, building search engines, or when neither approach alone provides sufficient recall. | [source](https://github.com/wshobson/agents/tree/main/plugins/llm-application-dev/skills/hybrid-search-implementation) |
| `rag-implementation` | adjacent | Build Retrieval-Augmented Generation (RAG) systems for LLM applications with vector databases and semantic search. Use when implementing knowledge-grounded AI, building document Q&A systems, or integrating LLMs with external knowledge … | [source](https://github.com/wshobson/agents/tree/main/plugins/llm-application-dev/skills/rag-implementation) |
## Source: ConardLi/garden-skills
- Repository: [https://github.com/ConardLi/garden-skills](https://github.com/ConardLi/garden-skills) (commit `fbd6453`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `kb-retriever` | adjacent | 面向本地知识库目录的检索和问答助手。核心流程:(1)分层索引导航 (2)遇到PDF/Excel时必须先读取references学习处理方法 (3)处理文件后再检索。按文件类型组合使用 grep、Read、pdfplumber、pandas 进行渐进式检索,避免整文件加载。用户问题涉及"从知识库目录回答问题/检索信息/查资料"时使用。 | [source](https://github.com/ConardLi/garden-skills/tree/fbd6453/skills/kb-retriever) |
## Source: a5c-ai/babysitter
- Repository: [https://github.com/a5c-ai/babysitter](https://github.com/a5c-ai/babysitter) (commit `44a5d58b`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `opencl-runtime` | core | Cross-vendor OpenCL runtime management and kernel development. Query platforms/devices, generate portable OpenCL C kernel code, handle vendor-specific extensions, manage contexts and command queues, compile and cache programs. | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/gpu-programming/skills/opencl-runtime) |
| `lms-admin` | core | Configure and manage Learning Management System operations including courses, enrollments, compliance training, and learning analytics | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/domains/business/human-resources/skills/lms-admin) |
| `CVE/CWE Database Skill` | core | CVE and CWE database querying and management | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/security-research/skills/cve-cwe-db) |
| `mlflow-experiment-tracker` | core | MLflow integration skill for experiment tracking, model registry, and artifact management. Enables LLMs to log experiments, compare runs, manage model lifecycle, and retrieve artifacts through the MLflow API. | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/data-science-ml/skills/mlflow-experiment-tracker) |
| `security-sandbox` | core | Isolated analysis environment management for malware and exploit testing. Create and manage isolated VMs, configure Cuckoo Sandbox, set up REMnux/FlareVM environments, manage Docker-based analysis containers, and capture filesystem and … | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/security-research/skills/security-sandbox) |
| `translation-management` | core | Integration with translation management systems and i18n workflows. Connect with Crowdin, Transifex, Weblate, manage translation memory, synchronize glossaries, and automate localization pipelines. | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/technical-documentation/skills/translation-management) |
| `unity-development` | core | Unity Engine integration skill for project setup, C# scripting, scene management, prefab creation, and editor automation. Enables LLMs to interact with Unity Editor through MCP servers for asset manipulation, script generation, and … | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/game-development/skills/unity-development) |
| `godot-development` | core | Godot Engine integration skill for GDScript/C# development, scene composition, node management, and editor automation. Enables LLMs to interact with Godot Editor through MCP servers for asset manipulation, script generation, and automated … | [source](https://github.com/a5c-ai/babysitter/tree/44a5d58b/library/specializations/game-development/skills/godot-development) |
## Source: brycewang-stanford/Auto-Empirical-Research-Skills
- Repository: [https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills) (commit `85bf545`, retrieved 2026-07-14)
- License: CC-BY-4.0
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `hud` | core | Diverga HUD (Heads-Up Display) management skill. Configure and manage the research project statusline display. Supports multiple presets: research, checkpoint, memory, minimal. Triggers: "hud", "statusline", "display settings | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/25-HosungYou-Diverga/skills/hud) |
| `humanities-skills` | core | 5 humanities skills. Trigger: textual analysis, archival research, digital humanities, philosophy. Design: digital tools and qualitative methods for humanities scholarship. | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/43-wentorai-research-plugins/skills/domains/humanities) |
| `search-skills` | core | 31 database search skills. Trigger: finding papers, search strategies, querying academic databases. Design: one skill per database/tool with API details, query syntax, and rate limits. | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/43-wentorai-research-plugins/skills/literature/search) |
| `bibliography-management-guide` | core | Manage references with BibLaTeX, natbib, and LaTeX bibliography styles | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/43-wentorai-research-plugins/skills/writing/latex/bibliography-management-guide) |
| `bibtex-management-guide` | core | Clean, format, deduplicate, and manage BibTeX bibliography files for LaTeX | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/43-wentorai-research-plugins/skills/writing/citation/bibtex-management-guide) |
| `i3` | core | RAG Builder with Parallel Document Processing Vector database construction with local embeddings (zero cost) Handles PDF download, text extraction, chunking, and vector database creation Absorbed B5 (Parallel Document Processor) … | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/25-HosungYou-Diverga/skills/i3) |
| `automation-skills` | core | 10 research automation skills. Trigger: automating experiments, tracking results, reproducible pipelines. Design: ML experiment management, workflow orchestration, and lab automation tools. | [source](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/tree/85bf545/skills/43-wentorai-research-plugins/skills/research/automation) |
## Source: davepoon/buildwithclaude
- Repository: [https://github.com/davepoon/buildwithclaude](https://github.com/davepoon/buildwithclaude) (commit `3c94e0c`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `supabase-automation` | core | Automate Supabase database queries, table management, project administration, storage, edge functions, and SQL execution via Rube MCP (Composio). Always search tools first for current schemas. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/supabase-automation) |
| `wrike-automation` | core | Automate Wrike project management via Rube MCP (Composio): create tasks/folders, manage projects, assign work, and track progress. Always search tools first for current schemas. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/wrike-automation) |
| `microsoft-teams-automation` | core | Automate Microsoft Teams tasks via Rube MCP (Composio): send messages, manage channels, create meetings, handle chats, and search messages. Always search tools first for current schemas. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/microsoft-teams-automation) |
| `slack-automation` | core | Automate Slack messaging, channel management, search, reactions, and threads via Rube MCP (Composio). Send messages, search conversations, manage channels/users, and react to messages programmatically. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/slack-automation) |
| `google-calendar-automation` | core | Automate Google Calendar events, scheduling, availability checks, and attendee management via Rube MCP (Composio). Create events, find free slots, manage attendees, and list calendars programmatically. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/google-calendar-automation) |
| `datadog-automation` | core | Automate Datadog tasks via Rube MCP (Composio): query metrics, search logs, manage monitors/dashboards, create events and downtimes. Always search tools first for current schemas. | [source](https://github.com/davepoon/buildwithclaude/tree/3c94e0c/plugins/all-skills/skills/datadog-automation) |
## Source: davila7/claude-code-templates
- Repository: [https://github.com/davila7/claude-code-templates](https://github.com/davila7/claude-code-templates) (commit `fa79251`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `omero-integration` | core | Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows. | [source](https://github.com/davila7/claude-code-templates/tree/fa79251/cli-tool/components/skills/scientific/omero-integration) |
| `openalex-database` | core | Query and analyze scholarly literature using the OpenAlex database. This skill should be used when searching for academic papers, analyzing research trends, finding works by authors or institutions, tracking citations, discovering open … | [source](https://github.com/davila7/claude-code-templates/tree/fa79251/cli-tool/components/skills/scientific/openalex-database) |
| `fda-database` | core | Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research. | [source](https://github.com/davila7/claude-code-templates/tree/fa79251/cli-tool/components/skills/scientific/fda-database) |
| `metabolomics-workbench-database` | core | Access NIH Metabolomics Workbench via REST API (4,200+ studies). Query metabolites, RefMet nomenclature, MS/NMR data, m/z searches, study metadata, for metabolomics and biomarker discovery. | [source](https://github.com/davila7/claude-code-templates/tree/fa79251/cli-tool/components/skills/scientific/metabolomics-workbench-database) |
## Source: foryourhealth111-pixel/Vibe-Skills
- Repository: [https://github.com/foryourhealth111-pixel/Vibe-Skills](https://github.com/foryourhealth111-pixel/Vibe-Skills) (commit `34429a8`, retrieved 2026-07-14)
- License: Apache-2.0
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `omero-integration` | core | Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows. | [source](https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/34429a8/bundled/skills/omero-integration) |
| `openalex-database` | core | Query and analyze scholarly literature using the OpenAlex database. This skill should be used when searching for academic papers, analyzing research trends, finding works by authors or institutions, tracking citations, discovering open … | [source](https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/34429a8/bundled/skills/openalex-database) |
| `fda-database` | core | Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research. | [source](https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/34429a8/bundled/skills/fda-database) |
| `metabolomics-workbench-database` | core | Access NIH Metabolomics Workbench via REST API (4,200+ studies). Query metabolites, RefMet nomenclature, MS/NMR data, m/z searches, study metadata, for metabolomics and biomarker discovery. | [source](https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/34429a8/bundled/skills/metabolomics-workbench-database) |
## Source: giuseppe-trisciuoglio/developer-kit
- Repository: [https://github.com/giuseppe-trisciuoglio/developer-kit](https://github.com/giuseppe-trisciuoglio/developer-kit) (commit `306f428`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `qdrant` | core | Provides Qdrant vector database integration patterns with LangChain4j. Handles embedding storage, similarity search, and vector management for Java applications. Use when implementing vector-based retrieval for RAG systems, semantic … | [source](https://github.com/giuseppe-trisciuoglio/developer-kit/tree/306f428/plugins/developer-kit-java/skills/qdrant) |
| `memory-md-management` | core | Provides comprehensive memory file management capabilities including auditing, quality assessment, and targeted improvements for files such as CLAUDE.md. Use when user asks to check, audit, update, improve, fix, maintain, or validate … | [source](https://github.com/giuseppe-trisciuoglio/developer-kit/tree/306f428/plugins/developer-kit-core/skills/memory-md-management) |
## Source: jawwadfirdousi/agent-skills
- Repository: [https://github.com/jawwadfirdousi/agent-skills](https://github.com/jawwadfirdousi/agent-skills) (commit `962ffed`, retrieved 2026-07-14)
- License: no explicit license — referenced by link only
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `supabase` | core | Run Supabase Management API SQL for persistent data tasks such as querying records, applying schema changes, managing policies, and handling storage metadata. Use when requests involve Supabase database CRUD, migrations, or production-like … | [source](https://github.com/jawwadfirdousi/agent-skills/tree/962ffed/supabase/skills/supabase) |
## Source: jeremylongshore/claude-code-plugins-plus-skills
- Repository: [https://github.com/jeremylongshore/claude-code-plugins-plus-skills](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) (commit `e112938a`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `windsurf-cascade-context` | core | Manage Cascade context window and memory for complex projects. Activate when users mention "cascade context", "ai memory", "context management", "large codebase navigation", or "multi-session development". Handles context optimization and … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/skill-databases/windsurf/skills/windsurf-cascade-context) |
| `shopify-core-workflow-b` | core | Manage Shopify orders, customers, and fulfillments using the GraphQL Admin API. Use when querying orders, processing fulfillments, managing customers, or building order management integrations. Trigger with phrases like "shopify orders", … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/shopify-pack/skills/shopify-core-workflow-b) |
| `notion-load-scale` | core | High-volume Notion operations: parallel requests within 3 req/sec, worker queues, database pagination at scale, incremental sync for large workspaces, and memory management for bulk operations. Trigger with phrases like "notion scale", … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/notion-pack/skills/notion-load-scale) |
| `windsurf-custom-prompts` | core | Create and manage custom prompt libraries for Cascade. Activate when users mention "custom prompts", "prompt library", "prompt templates", "cascade prompts", or "prompt management". Handles prompt library creation and organization. Use … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/skill-databases/windsurf/skills/windsurf-custom-prompts) |
| `intercom-core-workflow-a` | core | Manage Intercom contacts: create, search, update, merge leads into users. Use when building contact management features, syncing user data, or implementing contact search and segmentation. Trigger with phrases like "intercom contacts", … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/intercom-pack/skills/intercom-core-workflow-a) |
| `anth-security-basics` | core | Apply Anthropic Claude API security best practices for key management, input validation, and prompt injection defense. Use when securing API keys, validating user inputs before sending to Claude, or implementing content safety guardrails. … | [source](https://github.com/jeremylongshore/claude-code-plugins-plus-skills/tree/e112938a/plugins/saas-packs/anthropic-pack/skills/anth-security-basics) |
## Source: K-Dense-AI/claude-scientific-skills
- Repository: [https://github.com/K-Dense-AI/claude-scientific-skills](https://github.com/K-Dense-AI/claude-scientific-skills) (commit `4d97e29`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `omero-integration` | core | Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows. | [source](https://github.com/K-Dense-AI/claude-scientific-skills/tree/4d97e29/skills/omero-integration) |
## Source: K-Dense-AI/scientific-agent-skills
- Repository: [https://github.com/K-Dense-AI/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills) (commit `4d97e29`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `omero-integration` | core | Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows. | [source](https://github.com/K-Dense-AI/scientific-agent-skills/tree/4d97e29/skills/omero-integration) |
## Source: nowork-studio/NotFair
- Repository: [https://github.com/nowork-studio/NotFair](https://github.com/nowork-studio/NotFair) (commit `99bda4b`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `google-ads` | core | Manage Google Ads — performance, keywords, bids, budgets, negatives, campaigns, ads, search terms, QS, location targeting, bulk operations, experiments, asset management, portfolio bidding, offline conversions. Use for any mention of … | [source](https://github.com/nowork-studio/NotFair/tree/99bda4b/google-ads/manage) |
## Source: ruvnet/claude-code-flow
- Repository: [https://github.com/ruvnet/claude-code-flow](https://github.com/ruvnet/claude-code-flow) (commit `73914bd`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `AgentDB Advanced Features` | core | Master advanced AgentDB features including QUIC synchronization, multi-database management, custom distance metrics, hybrid search, and distributed systems integration. Use when building distributed AI systems, multi-agent coordination, or … | [source](https://github.com/ruvnet/claude-code-flow/tree/73914bd/.agents/skills/agentdb-advanced) |
## Source: ruvnet/ruflo
- Repository: [https://github.com/ruvnet/ruflo](https://github.com/ruvnet/ruflo) (commit `73914bd`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `AgentDB Advanced Features` | core | Master advanced AgentDB features including QUIC synchronization, multi-database management, custom distance metrics, hybrid search, and distributed systems integration. Use when building distributed AI systems, multi-agent coordination, or … | [source](https://github.com/ruvnet/ruflo/tree/73914bd/.agents/skills/agentdb-advanced) |
## Source: samber/cc-skills-golang
- Repository: [https://github.com/samber/cc-skills-golang](https://github.com/samber/cc-skills-golang) (commit `4881c01`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `golang-security` | core | Security best practices and vulnerability prevention for Golang. Covers injection (SQL, command, XSS), cryptography, filesystem safety, network security, cookies, secrets management, memory safety, and logging. Apply when writing, … | [source](https://github.com/samber/cc-skills-golang/tree/4881c01/skills/golang-security) |
## Source: secondsky/claude-skills
- Repository: [https://github.com/secondsky/claude-skills](https://github.com/secondsky/claude-skills) (commit `c4889f6`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `mcp-management` | core | Manage MCP servers - discover, analyze, execute tools/prompts/resources. Use for MCP integrations, capability discovery, tool filtering, programmatic execution, or encountering context bloat, server configuration, tool execution errors. | [source](https://github.com/secondsky/claude-skills/tree/c4889f6/plugins/mcp-management/skills/mcp-management) |
## Source: sgcarstrends/backend
- Repository: [https://github.com/sgcarstrends/backend](https://github.com/sgcarstrends/backend) (commit `7231cbc`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `logo-management` | core | Manage car logo fetching, scraping, and Vercel Blob storage in the logos package. Use when adding new car brand logos, updating logo sources, debugging brand name normalization, managing Vercel Blob storage, or optimizing logo caching. | [source](https://github.com/sgcarstrends/backend/tree/7231cbc/.agents/skills/logo-management) |
## Source: wondelai/skills
- Repository: [https://github.com/wondelai/skills](https://github.com/wondelai/skills) (commit `326b380`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `high-output-management` | core | Manage for output using Grove''s "High Output Management": a manager''s output is their organization''s output, raised by high-leverage activities. Use when the user mentions "high output management", "managerial leverage", "one-on-ones", … | [source](https://github.com/wondelai/skills/tree/326b380/high-output-management) |
## Source: NVIDIA/skills
- Repository: [https://github.com/NVIDIA/skills](https://github.com/NVIDIA/skills) (commit `153b14b`, retrieved 2026-07-14)
- License: CC-BY-4.0 (skills/docs), Apache-2.0 (code)
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `nemo-rl-session-memory` | core | Manage durable working-session memory for coding agents. Use when a user asks to preserve or recover agent context across disconnects, VS Code restarts, long-running work, handoffs, or any session where important state should be written … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-rl-session-memory) |
| `vss-manage-video-io-storage` | core | Use to call the VIOS REST API (sensor list, timelines, clip extraction, snapshots, add/delete sensors and streams). Not for VLM inference or search. | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/vss-manage-video-io-storage) |
| `rag-blueprint` | core | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. Handles any RAG action: deploy, install, start, enable, disable, toggle, change, configure, troubleshoot, debug, fix, shutdown, stop, or tear down any RAG feature or … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/rag-blueprint) |
| `nemo-rl-docs` | core | Documentation conventions for NeMo-RL. Covers docs/index.md updates and docstring format. Do NOT use for: bug fixes, test fixes, dependency bumps, refactoring, CI/CD changes, performance tuning, or any task that does not involve writing or … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-rl-docs) |
| `nemoclaw-user-guide` | core | Guides human users' AI agents to the NemoClaw docs MCP server and canonical Fern documentation in Markdown form. Use when users ask how to install, configure, operate, troubleshoot, secure, or learn NemoClaw with an AI coding assistant. … | [source](https://github.com/NVIDIA/skills/tree/153b14b/plugins/nvidia-skills/skills/nemoclaw-user-guide) |
| `tilegym-cutile-python` | core | Expert cuTile programming assistant. Write high-performance GPU kernels using cuTile's tile-based programming model with proper validation and optimization. Supports deep agent orchestration for complex multi-kernel tasks. | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/tilegym-cutile-python) |
| `earth2studio-discover` | core | Find Earth2Studio models, data sources, and examples for a weather/climate use case. Do NOT use for writing inference code, downloading data, or installation. | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/earth2studio-discover) |
| `earth2studio-install` | core | Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. Do NOT use for writing inference code, choosing models, or PhysicsNeMo questions. | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/earth2studio-install) |
| `nemo-automodel-model-onboarding` | core | Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation. | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-automodel-model-onboarding) |
| `nemo-mbridge-multi-node-slurm` | core | Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-mbridge-multi-node-slurm) |
| `nemo-rl-auto-research` | core | Autonomous NeMo-RL research agent workflow for directed hypothesis testing and open-ended discovery. Guides agents through the full experiment lifecycle: understanding recipes and environments, wiring RL or NeMo-gym runs, launching … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-rl-auto-research) |
| `nemo-rl-brev-etiquette` | core | Brev instance operating guidance for NeMo-RL agents working in /home/ubuntu/RL with limited workspace disk, a larger /ephemeral volume, and optional /home/ubuntu/RL/.env secrets. Use when running nemo-rl-auto-research campaigns, … | [source](https://github.com/NVIDIA/skills/tree/153b14b/skills/nemo-rl-brev-etiquette) |
## Source: veniceai/skills
- Repository: [https://github.com/veniceai/skills](https://github.com/veniceai/skills) (commit `de089fa`, retrieved 2026-07-14)
- License: MIT
| Skill | Tier | What it adds | Upstream |
|---|---|---|---|
| `venice-api-keys` | core | Manage Venice API keys. Covers GET/POST/PATCH/DELETE /api_keys, GET /api_keys/{id}, GET /api_keys/rate_limits, GET /api_keys/rate_limits/log, the two-step /api_keys/generate_web3_key wallet flow, INFERENCE vs ADMIN key types, and per-key … | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-api-keys) |
| `venice-characters` | adjacent | Discover and use Venice public characters (persona-driven system prompts with a bound model). Covers GET /characters (search/filter/sort), /characters/{slug}, /characters/{slug}/reviews, the Character schema, and how to apply a character … | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-characters) |
| `venice-image-edit` | adjacent | Transform existing images with Venice. Covers POST /image/edit (prompt-driven single-image edit), /image/multi-edit (compose 1-3 images), /image/upscale (2-4x upscale + enhance), and /image/background-remove. Accepts base64, file upload, … | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-image-edit) |
| `venice-image-generate` | adjacent | Generate images with Venice. Covers POST /image/generate (Venice-native), POST /images/generations (OpenAI-compatible), GET /image/styles (style presets), request fields (prompt, dimensions, cfg_scale, seed, variants, style_preset, … | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-image-generate) |
| `venice-embeddings` | adjacent | Call POST /embeddings on Venice. Covers request shape (input, model, encoding_format, dimensions, user), OpenAI compatibility, response compression (gzip/br), and practical usage for retrieval, clustering, and RAG. | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-embeddings) |
| `venice-errors` | adjacent | Handle Venice API errors correctly. Covers the StandardError / DetailedError / ContentViolationError / X402InferencePaymentRequired body shapes, every meaningful status code (400, 401, 402, 403, 415, 422, 429, 500, 503, 504), the 402 … | [source](https://github.com/veniceai/skills/tree/de089fa/skills/venice-errors) |

58
references/market.md Normal file
View File

@@ -0,0 +1,58 @@
# Market evidence report — big-data-archive-librarian
Source: **15 real job ads** (JSearch API, countries: us 15), extracted into the MSSQL evidence store; as of 2026-07-09.
This report contains extracted, aggregated facts only — no ad text is
reproduced (copyright / platform terms).
## Seniority distribution
| Seniority | Ads | Share |
|---|---|---|
| mid | 11 | 73 % |
| junior | 2 | 13 % |
| n/a | 1 | 7 % |
| senior | 1 | 7 % |
## Tools — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | ArchivesSpace | 4 | 27 % |
| 2 | Microsoft Office | 3 | 20 % |
## Hard skills — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | metadata creation | 6 | 40 % |
| 2 | digital preservation | 5 | 33 % |
| 3 | archival arrangement | 3 | 20 % |
| 4 | cataloging | 3 | 20 % |
| 5 | digitization | 3 | 20 % |
## Methods — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
## Responsibilities — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
## Job title variants in the market
| Title | Ads |
|---|---|
| Digital Archivist | 6 |
| Archival Metadata Librarian — Physical Archives | 1 |
| Archives Technician | 1 |
| Archivist - Library | 1 |
| Assistant Librarian: Digital Processing Archivist | 1 |
| Digital Archivist Intern | 1 |
| Historian (Vacancy#:VAR003385) | 1 |
| Librarian/Cataloger | 1 |
| Librarian/Cataloger | 1 |
| Library Technician (Metadata and Digitization Support) | 1 |
Methodology: entities extracted per ad ({hard_skills, tools, methods, responsibilities, seniority}), normalized, counted as DISTINCT ads per entity; report threshold ≥ 3 ads. Headline sections in skills.md/tools.md use the stricter ≥ 20 % threshold.

22
references/profile.md Normal file
View File

@@ -0,0 +1,22 @@
# Occupation profile — big data archive librarian
- **ESCO URI:** http://data.europa.eu/esco/occupation/4e4a1aa1-ca53-406e-9ca6-889d4621a0ed
- **ESCO code:** 3433.2
- **ISCO-08 group:** 3433 — Gallery, museum and library technicians
- **O*NET-SOC:** 43-4121.00 — Library Assistants, Clerical (match: closeMatch)
## Description (ESCO)
Big data archive librarians classify, catalogue and maintain libraries of digital media. They also evaluate and comply with metadata standards for digital content and update obsolete data and legacy systems.
## Definition
nan
## Alternative labels
- digital documentation archivist
- archive librarian
- documentation archivist
- computer tape librarian
- digital archivists

98
references/skills.md Normal file
View File

@@ -0,0 +1,98 @@
# Competences — big data archive librarian
Source: ESCO v1.2.1 occupation-skill relations (http://data.europa.eu/esco/occupation/4e4a1aa1-ca53-406e-9ca6-889d4621a0ed).
## Essential
- **analyse big data** (skill/competence)
- **business intelligence** (knowledge)
- **comply with legal regulations** (skill/competence)
- **data extraction, transformation and loading tools** (knowledge)
- **data models** (knowledge)
- **database** (knowledge)
- **database development tools** (knowledge)
- **database management systems** (knowledge)
- **digital curation** (knowledge)
- **digital data processing** (nan)
- **digitization** (knowledge)
- **maintain data entry requirements** (skill/competence)
- **maintain database performance** (skill/competence)
- **maintain database security** (skill/competence)
- **manage archive users guidelines** (skill/competence)
- **manage content metadata** (skill/competence)
- **manage data** (skill/competence)
- **manage database** (skill/competence)
- **manage digital archives** (skill/competence)
- **manage ICT data classification** (skill/competence)
- **query languages** (knowledge)
- **resource description framework query language** (knowledge)
- **write database documentation** (skill/competence)
## Optional
- apply information security policies (skill/competence)
- CA Datacom/DB (knowledge)
- creatively use digital technologies (skill/competence)
- data quality assessment (knowledge)
- DB2 (knowledge)
- design database backup specifications (skill/competence)
- design database scheme (skill/competence)
- develop ICT workflow (skill/competence)
- digitise documents (skill/competence)
- Filemaker (database management systems) (knowledge)
- give live presentation (skill/competence)
- IBM Informix (knowledge)
- IBM InfoSphere DataStage (knowledge)
- IBM InfoSphere Information Server (knowledge)
- identify technological needs (skill/competence)
- Informatica PowerCenter (knowledge)
- information confidentiality (knowledge)
- information structure (knowledge)
- integrate ICT data (skill/competence)
- LDAP (knowledge)
- LINQ (knowledge)
- manage data collection systems (skill/competence)
- manage information access aids (skill/competence)
- MarkLogic (knowledge)
- MDX (knowledge)
- Microsoft Access (knowledge)
- migrate existing data (skill/competence)
- monitor technology trends (skill/competence)
- MySQL (knowledge)
- N1QL (knowledge)
- normalise data (skill/competence)
- ObjectStore (knowledge)
- OpenEdge Database (knowledge)
- Oracle Data Integrator (knowledge)
- Oracle Relational Database (knowledge)
- Oracle Warehouse Builder (knowledge)
- Pentaho Data Integration (knowledge)
- perform backups (skill/competence)
- PostgreSQL (knowledge)
- QlikView Expressor (knowledge)
- SAP Data Services (knowledge)
- SAS Data Management (knowledge)
- SPARQL (knowledge)
- SQL Server (knowledge)
- SQL Server Integration Services (knowledge)
- statistics (knowledge)
- Teradata Database (knowledge)
- TripleStore (knowledge)
- visual presentation techniques (knowledge)
- XQuery (knowledge)
<!-- market-evidence -->
## Market evidence (job-ad analysis, 15 ads, as of 2026-07-09)
Share of analyzed job ads mentioning the item (threshold ≥ 20 %). Source: JSearch/Adzuna APIs.
### Hard skills
- metadata creation — **40 %**
- digital preservation — **33 %**
- digitization — **20 %**
- archival arrangement — **20 %**
- cataloging — **20 %**
<!-- market-evidence -->

77
references/tasks.md Normal file
View File

@@ -0,0 +1,77 @@
# Tasks & work activities — big data archive librarian
Source: O*NET 30.3, occupation 43-4121.00 (Library Assistants, Clerical).
## Task statements
- **[Core]** Open and close library during specified hours and secure library equipment, such as computers and audio-visual equipment.
- **[Core]** Perform clerical activities, such as answering phones, sorting mail, filing, typing, word processing, and photocopying and mailing out material.
- **[Core]** Schedule, supervise, and train clerical workers, volunteers, student assistants, and other library employees.
- **[Core]** Maintain library equipment, such as photocopiers, scanners, and computers, and instruct patrons in proper use of such equipment.
- **[Core]** Manage reserve materials by placing items on reserve for library patrons, checking items in and out of library, and removing out-of-date items.
- **[Core]** Lend, reserve, and collect books, periodicals, videotapes, and other materials at circulation desks and process materials for inter-library loans.
- **[Core]** Repair books using mending tape, paste, and brushes or prepare books to be sent to a bindery for repair.
- **[Core]** Prepare library statistics reports.
- **[Core]** Enter and update patrons' records on computers.
- **[Core]** Process new materials including books, audio-visual materials, and computer software.
- **[Core]** Sort books, publications, and other items according to established procedure and return them to shelves, files, or other designated storage areas.
- **[Core]** Locate library materials for patrons, including books, periodicals, tape cassettes, Braille volumes, and pictures.
- **[Core]** Instruct patrons on how to use reference sources, card catalogs, and automated information systems.
- **[Core]** Inspect returned books for condition and due-date status and compute any applicable fines.
- **[Core]** Answer routine inquiries and refer patrons in need of professional assistance to librarians.
- **[Core]** Maintain records of items received, stored, issued, and returned and file catalog cards according to system used.
- **[Core]** Provide assistance to librarians in the maintenance of collections of books, periodicals, magazines, newspapers, and audio-visual and other materials.
- **[Core]** Take action to deal with disruptive or problem patrons.
- **[Core]** Register new patrons and issue borrower identification cards that permit patrons to borrow books and other materials.
- **[Core]** Send out notices and accept fine payments for lost or overdue books.
- **[Core]** Prepare, store, and retrieve classification and catalog information, lecture notes, or other information related to stored documents, using computers.
- **[Core]** Review records, such as microfilm and issue cards, to identify titles of overdue materials and delinquent borrowers.
- **[Core]** Select substitute titles when requested materials are unavailable, following criteria such as age, education, and interests.
- **[Core]** Deliver and retrieve items to and from departments by hand or using push carts.
- **[Core]** Assist in the preparation of book displays.
- **[Supplemental]** Perform accounting and bookkeeping activities, such as invoicing, maintaining financial records, budgeting, and handling cash.
- **[Supplemental]** Acquire books, pamphlets, periodicals, audio-visual materials, and other library supplies by checking prices, figuring costs, and preparing appropriate order forms and facilitating the ordering process by providing such information to others.
- **[Supplemental]** Design or maintain library web site and online catalogues.
- **[Supplemental]** Plan or participate in library events and programs, such as story time with children.
- **[Supplemental]** Classify and catalog items according to content and purpose.
- **[Supplemental]** Operate small branch libraries, under the direction of off-site librarian supervisors.
- **[Supplemental]** Operate and maintain audio-visual equipment.
## Detailed work activities
- Answer telephones to direct calls or provide information.
- Arrange items for use or display.
- Calculate financial data.
- Collect deposits, payments or fees.
- Deliver items.
- Demonstrate activity techniques or equipment use.
- Develop computer or online applications.
- Distribute materials to employees or customers.
- Enter information into databases or software programs.
- Inspect items for damage or defects.
- Issue documentation or identification to customers or employees.
- Maintain electronic equipment.
- Maintain financial or account records.
- Maintain inventories of materials, equipment, or products.
- Maintain inventory records.
- Maintain office equipment in proper operating condition.
- Maintain security.
- Manage clerical or administrative activities.
- Operate office equipment.
- Order materials, supplies, or equipment.
- Plan educational activities.
- Plan special events.
- Prepare documentation for contracts, transactions, or regulatory compliance.
- Prepare employee work schedules.
- Prepare research or technical reports.
- Process library materials.
- Provide customer service to clients or users.
- Refer customers to appropriate personnel.
- Repair books or other printed material.
- Send information, materials or documentation.
- Sort mail.
- Sort materials or products.
- Store records or related materials.
- Supervise clerical or administrative personnel.
- Track goods or materials.
- Type documents.

39
references/tools.md Normal file
View File

@@ -0,0 +1,39 @@
# Tools & technology — big data archive librarian
Source: O*NET 30.3 'Software Skills' for 43-4121.00.
| Software | Category | Hot technology |
|---|---|---|
| Microsoft Access | Data base user interface and query software | yes |
| Adobe Acrobat | Document management software | yes |
| Microsoft Outlook | Electronic mail software | yes |
| Adobe Photoshop | Graphics or photo imaging software | yes |
| C++ | Object or component oriented development software | yes |
| Microsoft Office software | Office suite software | yes |
| Microsoft Windows | Operating system software | yes |
| Microsoft PowerPoint | Presentation software | yes |
| Microsoft Excel | Spreadsheet software | yes |
| Microsoft Word | Word processing software | yes |
| Database software | Data base user interface and query software | |
| Recordkeeping software | Data base user interface and query software | |
| Microsoft Publisher | Desktop publishing software | |
| Video retrieval systems | Information retrieval or search software | |
| Web browser software | Internet browser software | |
| Automated circulation systems | Library software | |
| Cataloging software | Library software | |
| Online Computer Library Center (OCLC) databases | Library software | |
| ResourceMate Plus | Library software | |
| WorldCat | Library software | |
<!-- market-evidence -->
## Market evidence (job-ad analysis, 15 ads, as of 2026-07-09)
Share of analyzed job ads mentioning the item (threshold ≥ 20 %). Source: JSearch/Adzuna APIs.
### Tools
- ArchivesSpace — **27 %**
- Microsoft Office — **20 %**
<!-- market-evidence -->