Files
2026-08-14 19:11:28 +02:00

38 KiB
Raw Permalink Blame History

External AI agent skills — sound-mastering-engineer

Proven, publicly available AI agent skills mapped to this occupation. Nothing is copied from the sources: every entry is a name, a one-line summary and a link to the upstream skill package. Each section names its source repository, commit, license and retrieval date.

Tiers: core = the skill directly exercises a top market hard skill, tool or method (from gated job-ad evidence) or an essential ESCO competence of this occupation; adjacent = plausibly useful, secondary. Entries are capped at 12 per source and 80 in total per occupation (core first, strongest matches survive); everything beyond the caps is excluded and logged in the pipeline audit trail, not in this package.

Matched deterministically (ISCO group + title/competence keywords, tiered against market evidence + ESCO essentials) by pipeline/p5_enrich_ai_skills.py on 2026-07-14.

Source: anthropics/skills

  • Repository: https://github.com/anthropics/skills (commit f6656c1, retrieved 2026-07-14)
  • License: Apache-2.0; the document skills (docx/pdf/pptx/xlsx) are source-available — see the LICENSE.txt in the upstream skill folder
Skill Tier What it adds Upstream
docx adjacent Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to … source
pdf adjacent Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating … source

Source: ConardLi/garden-skills

Skill Tier What it adds Upstream
web-video-presentation adjacent 把一篇文章或口播稿,做成"看起来像视频"的点击驱动 16:9 网页演示,可选合成口播音频。流程:原始文章 → 一次产出口播稿 + outline 开发计划 → 用户一次对齐 5 件事(稿子 / outline / 主题 / 素材 / 开发模式)→ 网页开发(逐章 / 顺序 / 并行)→ 可选音频合成provider-agnostic内置 MiniMax mmx-cli + OpenAI TTS可换 ElevenLabs / edge-tts / … source

Source: a5c-ai/babysitter

Skill Tier What it adds Upstream
sound-design-direction core Create comprehensive audio design including music cues, sound effects, Foley, and score direction source
procedural-audio core Procedural sound skill for synthesis and dynamic sound design. source
audio-dsp core Audio DSP skill for filters and real-time processing. source
music-prompt-engineering core Optimize and format prompts specifically for AI music generation platforms like Suno and Udio, including platform-specific syntax and tag optimization source
lyric-writing core Write complete song lyrics with structural annotations and production notes optimized for AI music generation platforms like Suno and Udio source
multimedia-learning-design adjacent Apply Mayer's multimedia learning principles to design effective audio, video, graphics, and animations that reduce cognitive load source

Source: anbeime/skill

Skill Tier What it adds Upstream
poetry-music-visual core 为古诗词提供配图与配乐的全流程创作指导支持深度解析诗词意境、生成画面描述、提供配乐创作蓝图Suno格式适用于诗词可视化、MV创作、文化传播等场景 source
video-creation-suite core 完整的视频创作套件支持原创创作、视频二创、视频分析三种模式集成Coze Bot API、Edge-TTS、Suno API涵盖多智能体协同、素材生成、视频合成全流程 source
dream-video-prompt-generator core 小省导购员数字人带货版即梦视频提示词生成系统,基于四大智能体协同(提示词生成师、质量管控师、知识库运维师、跨环节适配师),按照"主体+运动+场景+(镜头语言+光影+氛围)"公式输出中英文双版提示词适配5s短视频。确保人物一致性、视觉连贯性、情绪连贯性支持知识库智能复用和跨工具适配Suno音乐、AI绘画为数字人带货视频提供高质量提示词生成服务。 source

Source: bitwize-music-studio/claude-ai-music-skills

Skill Tier What it adds Upstream
sheet-music-publisher core Converts mastered audio to sheet music and creates printable songbooks. Use after mastering when the user wants sheet music or a songbook for their album. source
mastering-engineer core Guides audio mastering for streaming platforms including loudness optimization and tonal balance. Use when the user has approved tracks and wants to master audio files. source
mix-engineer core Polishes raw Suno audio by processing per-stem WAVs (vocals, backing_vocals, drums, bass, guitar, keyboard, strings, brass, woodwinds, percussion, synth, other) with targeted cleanup, EQ, and compression, then remixing into a polished … source
promo-director core Generates 15-second vertical promo videos for social media from mastered audio. Use after mastering is complete and before release, when the user wants social media content. source
import-audio core Moves audio files to the correct album location with proper path structure. Use when the user has downloaded WAV files from Suno or other sources that need to be organized. source

Source: brycewang-stanford/Auto-Empirical-Research-Skills

Skill Tier What it adds Upstream
markitdown core Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing … source

Source: davepoon/buildwithclaude

Skill Tier What it adds Upstream
tubeify adjacent Remove pauses, filler words (um, uh), and dead air from raw YouTube recordings via the Tubeify API. Use when the user wants to edit a video, clean up audio, trim silences, or polish a raw recording for YouTube. source

Source: davila7/claude-code-templates

Skill Tier What it adds Upstream
game-audio core Game audio principles. Sound design, music integration, adaptive audio systems. source
transformers core This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, … source
video-downloader core Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options. source
motion-canvas core Complete production-ready guide for Motion Canvas with ESM/CommonJS workarounds, full setup templates, and troubleshooting for programmatic video creation using TypeScript source
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source
nemo-curator adjacent GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source
markitdown adjacent Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more. source

Source: degausai/wonda

Skill Tier What it adds Upstream
wonda-cli adjacent Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation source

Source: eduardo-sl/go-agent-skills

Skill Tier What it adds Upstream
go-troubleshooting core Diagnose runtime problems in Go programs: panics and stack traces, deadlocks, goroutine leaks, memory leaks, OOM kills, race reports, and debugging with delve and pprof. Use when: "debug this panic", "read this stack trace", "deadlock", … source

Source: EveryInc/compound-engineering-plugin

Skill Tier What it adds Upstream
ce-riffrec-feedback-analysis adjacent Analyze Riffrec feedback captures from bundles or standalone recordings. Always load for riffrec-*.zip, session.json + events.json + recording.webm + voice.webm bundles, .mp4/.mov/.webm videos, .m4a/.mp3/.wav audio, … source

Source: foryourhealth111-pixel/Vibe-Skills

Skill Tier What it adds Upstream
transformers core This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, … source
xan core High-performance CSV processing with xan CLI for large tabular datasets, streaming transformations, and low-memory pipelines. source
markitdown adjacent Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more. source

Source: giuseppe-trisciuoglio/developer-kit

Skill Tier What it adds Upstream
memory-md-management core Provides comprehensive memory file management capabilities including auditing, quality assessment, and targeted improvements for files such as CLAUDE.md. Use when user asks to check, audit, update, improve, fix, maintain, or validate … source

Source: gooseworks-ai/goose-skills

Skill Tier What it adds Upstream
mix-master adjacent Canonical short-form-ad audio mix in one FFmpeg pass. VO loudnorm + 3.0× per-clip + 2.0× mix, music 0.13 base + apad+afade, sidechain compress 20:1 @ 0.01, climax line +20%, optional video-to-music duration sync. Replaces the reactive … source
create-video-seedance-2-fal adjacent Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio via generate_audio: true, internal multi-cut … source

Source: indranilbanerjee/digital-marketing-pro

Skill Tier What it adds Upstream
c2pa-metadata adjacent Embed C2PA (Content Authenticity Initiative) provenance manifests in AI-generated marketing assets (image/video/audio/PDF). Use when: preparing AI-generated ad creative, social images, or video for EU markets to comply with EU AI Act … source

Source: jeremylongshore/claude-code-plugins-plus-skills

Skill Tier What it adds Upstream
granola-performance-tuning core Optimize Granola transcription accuracy, note quality, and processing speed. Use when improving transcription quality, reducing processing time, optimizing templates for better AI output, or tuning audio setup. Trigger: "granola … source
deepgram-core-workflow-a core Implement production pre-recorded speech-to-text with Deepgram. Use when building audio transcription, batch processing, or implementing diarization and intelligence features. Trigger: "deepgram transcription", "speech to text", … source
speak-sdk-patterns core Production patterns for Speak language learning API: conversation sessions, pronunciation assessment, audio preprocessing, and batch operations. Use when implementing sdk patterns features, or troubleshooting Speak language learning … source
speak-prod-checklist core Production readiness checklist for Speak language learning integrations: auth, audio pipeline, monitoring, and compliance. Use when implementing prod checklist features, or troubleshooting Speak language learning integration issues. … source
abridge-performance-tuning core Optimize Abridge clinical AI integration performance for high-volume deployments. Use when reducing note generation latency, optimizing audio streaming throughput, improving FHIR push performance, or scaling for multi-site health systems. … source
deepgram-performance-tuning core Optimize Deepgram API performance for faster transcription and lower latency. Use when improving transcription speed, reducing latency, or optimizing audio processing pipelines. Trigger: "deepgram performance", "speed up deepgram", … source
speak-common-errors core Diagnose and fix common Speak API errors: authentication failures, audio format issues, rate limits, and session management problems. Use when implementing common errors features, or troubleshooting Speak language learning integration … source
speak-cost-tuning core Optimize Speak API costs through usage monitoring, tier selection, and efficient audio processing. Use when implementing cost tuning, or managing Speak language learning platform operations. Trigger with phrases like "speak cost tuning", … source
speak-debug-bundle core Collect diagnostic information for Speak API issues: auth verification, audio format validation, session inspection, and network testing. Use when implementing debug bundle features, or troubleshooting Speak language learning integration … source
speak-security-basics core Security best practices for Speak API keys, audio data privacy, student data protection, and COPPA/FERPA compliance. Use when implementing security basics features, or troubleshooting Speak language learning integration issues. Trigger … source
twinmind-security-basics core Security best practices for TwinMind: on-device audio processing, encrypted cloud backups, microphone permissions, and data privacy controls. Use when implementing security basics, or managing TwinMind meeting AI operations. Trigger with … source
elevenlabs-reference-architecture core Implement ElevenLabs reference architecture for production TTS/voice applications. Use when designing new ElevenLabs integrations, reviewing project structure, or building a scalable audio generation service. Trigger: "elevenlabs … source

Source: K-Dense-AI/claude-scientific-skills

Skill Tier What it adds Upstream
markitdown adjacent Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more. source

Source: K-Dense-AI/scientific-agent-skills

Skill Tier What it adds Upstream
markitdown adjacent Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more. source

Source: mbailey/voicemode

Skill Tier What it adds Upstream
voicemode core Voice interaction for Claude Code. Use when users mention voice mode, speak, talk, converse, voice status, or voice troubleshooting. source

Source: microsoft/skills

Skill Tier What it adds Upstream
azure-ai-contentunderstanding-py adjacent Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document … source
azure-ai-voicelive-py adjacent Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, … source

Source: mohitagw15856/pm-claude-skills

Skill Tier What it adds Upstream
youtube-script-writer core Write engaging, high-retention YouTube video scripts with visual and audio cues. Use when asked to write a YouTube script, design a video outline, draft a video hook, or structure a video narrative. Produces a polished script with multiple … source

Source: nexu-io/open-design

Skill Tier What it adds Upstream
audio-jingle core Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3, and SFX to ElevenLabs SFX or AudioCraft. Output is one MP3/WAV file … source
video-downloader core Download videos from YouTube and other platforms for offline viewing, editing, or archival with support for various formats and quality options. source
8-bit-orbit-video-template core Hyperframes-based video template for retro pixel deck motion design. Use when users want a high-fidelity, multi-scene HTML-to-video composition with advanced transitions, interactive preview controls, and ready-to-render default style. source
social-spotify-card core Spotify Now Playing-style card with album art, progress bar, and playback controls, suited to video overlays or personal homepages. source
venice-audio-music adjacent Music generation queueing, retrieval, and completion endpoints via Venice.ai. Suited for jingles, background loops, and prototype scoring. source
venice-audio-speech adjacent Text-to-speech models, voices, formats, and streaming via Venice.ai. Useful for narration, voiceover, and conversational agent voices. source

Source: NoizAI/skills

Skill Tier What it adds Upstream
sound-fx adjacent Use this skill whenever the user wants to generate sound effects, ambient audio, or short audio clips from a text description. Triggers include: any mention of 'sound effect', 'sfx', 'generate sound', 'make a sound', 'audio effect', … source
daily-news-caster adjacent Fetches the latest news using news-aggregator-skill, formats it into a podcast script in Markdown format, and uses the tts skill to generate a podcast audio file. Use when the user asks to get the latest news and read it out as a podcast. source

Source: OpenSenseNova/SenseNova-Skills

Skill Tier What it adds Upstream
outlier-detection-and-quality-assessment core 执行全面的异常值检测与数据质量评估,利用 IQR 方法识别异常值并结合偏度、峰度分析数据分布特征,适用于非正态分布数据的预处理阶段。 source

Source: openweb-org/openweb

Skill Tier What it adds Upstream
youtube-music core source

Source: Orchestra-Research/AI-research-SKILLs

Skill Tier What it adds Upstream
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source
nemo-curator adjacent GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source

Source: SamurAIGPT/Generative-Media-Skills

Skill Tier What it adds Upstream
muapi-media-generation core Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5 source
muapi-cinema-director core Direct high-fidelity cinematic video with AI — translates creative intent into technical cinematographic directives for Veo3, Kling, and Luma video models via muapi.ai source
muapi-seedance-2 core Expert Cinema Director skill for Seedance 2.0 (ByteDance) — high-fidelity video generation across Chinese, Global, and VIP tiers. Supports text-to-video, image-to-video, first-last-frame, omni reference, character training, omni-reference … source
muapi-ugc-video-factory adjacent Turn a person photo + a product photo + an optional script into a vertical 9:16 UGC-style video ad. Generates a lifestyle hero image (Nano-Banana Pro Edit), then animates it with native audio using Seedance 2.0 VIP image-to-video. source

Source: sanjay3290/ai-skills

Skill Tier What it adds Upstream
google-tts adjacent Convert documents and text to audio using Google Cloud Text-to-Speech. Use this skill when the user wants to: narrate a document, read aloud text, generate audio from a file, convert text to speech, create a recording of documentation or … source

Source: Vincentwei1021/video-shotcraft

Skill Tier What it adds Upstream
video-shotcraft core Create cinematic product videos from shot recipe cards, a validated template, and code/audio assets (Remotion + real page screenshots + 2.5D camera moves + beat-synced cuts + sound design). Use when the user asks to turn a frontend project … source

Source: zechenzhangAGI/AI-research-SKILLs

Skill Tier What it adds Upstream
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source
nemo-curator adjacent GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source

Source: google/skills

Skill Tier What it adds Upstream
ima-sdk-basics adjacent Use this skill for Interactive Media Ads (IMA) SDK client-side ad insertion when you are requesting video ads client-side into websites, apps, TVs or other platforms with VAST or VMAP. Do not use for Dynamic Ad Insertion (DAI), SSAI, or … source

Source: veniceai/skills

Skill Tier What it adds Upstream
venice-audio-speech core Generate speech from text via POST /audio/speech. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash), voices per family, output formats (mp3/opus/aac/flac/wav/pcm), streaming, … source
venice-audio-transcription core Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), supported formats (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text), timestamps, and … source
venice-chat core Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning … source
venice-audio-music adjacent Async music / audio-track generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental, voice selection, duration, language, speed, model capability probing, and … source
venice-video adjacent Generate and transcribe videos via Venice. Covers the async /video/quote + /video/queue + /video/retrieve + /video/complete loop, text-to-video, image-to-video, video-to-video (upscale), audio input, reference images, scene and element … source