Files
boom-operator/references/ai-skills.md
2026-08-14 19:11:44 +02:00

39 KiB
Raw Blame History

External AI agent skills — boom-operator

Proven, publicly available AI agent skills mapped to this occupation. Nothing is copied from the sources: every entry is a name, a one-line summary and a link to the upstream skill package. Each section names its source repository, commit, license and retrieval date.

Tiers: core = the skill directly exercises a top market hard skill, tool or method (from gated job-ad evidence) or an essential ESCO competence of this occupation; adjacent = plausibly useful, secondary. Entries are capped at 12 per source and 80 in total per occupation (core first, strongest matches survive); everything beyond the caps is excluded and logged in the pipeline audit trail, not in this package.

Matched deterministically (ISCO group + title/competence keywords, tiered against market evidence + ESCO essentials) by pipeline/p5_enrich_ai_skills.py on 2026-07-14.

Source: anthropics/skills

  • Repository: https://github.com/anthropics/skills (commit f6656c1, retrieved 2026-07-14)
  • License: Apache-2.0; the document skills (docx/pdf/pptx/xlsx) are source-available — see the LICENSE.txt in the upstream skill folder
Skill Tier What it adds Upstream
docx adjacent Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to … source
pdf adjacent Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating … source

Source: ConardLi/garden-skills

Skill Tier What it adds Upstream
web-video-presentation adjacent 把一篇文章或口播稿,做成"看起来像视频"的点击驱动 16:9 网页演示,可选合成口播音频。流程:原始文章 → 一次产出口播稿 + outline 开发计划 → 用户一次对齐 5 件事(稿子 / outline / 主题 / 素材 / 开发模式)→ 网页开发(逐章 / 顺序 / 并行)→ 可选音频合成provider-agnostic内置 MiniMax mmx-cli + OpenAI TTS可换 ElevenLabs / edge-tts / … source

Source: a5c-ai/babysitter

Skill Tier What it adds Upstream
sound-design-direction core Create comprehensive audio design including music cues, sound effects, Foley, and score direction source
procedural-audio adjacent Procedural sound skill for synthesis and dynamic sound design. source
ipa-transcription-phonological adjacent Transcribe speech using International Phonetic Alphabet and analyze sound systems including phonotactics and phonological rules source
wwise adjacent Wwise integration skill for sound banks, RTPC, and interactive music. source
audio-dsp adjacent Audio DSP skill for filters and real-time processing. source

Source: affaan-m/everything-claude-code

Skill Tier What it adds Upstream
fal-ai-media adjacent Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to … source

Source: AgriciDaniel/claude-blog

Skill Tier What it adds Upstream
blog-audio adjacent Generate audio narration of blog posts using Google Gemini TTS. Supports summary narration, full article read-aloud, and two-speaker podcast/dialogue mode with 30 voice options. Outputs MP3 with HTML5 audio embed code. Works standalone via … source

Source: bitwize-music-studio/claude-ai-music-skills

Skill Tier What it adds Upstream
promo-director adjacent Generates 15-second vertical promo videos for social media from mastered audio. Use after mastering is complete and before release, when the user wants social media content. source
import-audio adjacent Moves audio files to the correct album location with proper path structure. Use when the user has downloaded WAV files from Suno or other sources that need to be organized. source
import-art adjacent Places album art files in the correct audio and content directory locations. Use when the user has generated or downloaded album artwork that needs to be saved. source

Source: davepoon/buildwithclaude

Skill Tier What it adds Upstream
resemble-detect core Deepfake detection and media safety — detect AI-generated audio, images, video, and text, trace synthesis sources, apply watermarks, verify speaker identity, and analyze media intelligence using Resemble AI source
tubeify core Remove pauses, filler words (um, uh), and dead air from raw YouTube recordings via the Tubeify API. Use when the user wants to edit a video, clean up audio, trim silences, or polish a raw recording for YouTube. source
image-enhancer adjacent Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts. source

Source: davila7/claude-code-templates

Skill Tier What it adds Upstream
nemo-curator core GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source
game-audio adjacent Game audio principles. Sound design, music integration, adaptive audio systems. source
image-enhancer adjacent Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts. source

Source: disler/claude-code-hooks-multi-agent-observability

Skill Tier What it adds Upstream
Video Processor adjacent Process video files with audio extraction, format conversion (mp4, webm), and Whisper transcription. Use when user mentions video conversion, audio extraction, transcription, mp4, webm, ffmpeg, or whisper transcription. source

Source: EveryInc/compound-engineering-plugin

Skill Tier What it adds Upstream
ce-riffrec-feedback-analysis core Analyze Riffrec feedback captures from bundles or standalone recordings. Always load for riffrec-*.zip, session.json + events.json + recording.webm + voice.webm bundles, .mp4/.mov/.webm videos, .m4a/.mp3/.wav audio, … source

Source: gooseworks-ai/goose-skills

Skill Tier What it adds Upstream
video-polish core Takes an existing screen recording or demo video and adds professional zoom/pan effects synchronized to the narration. Uses transcript-driven zoom targeting and Remotion for rendering. Optionally replaces audio with a soundtrack. source
review-ugc-render adjacent Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates set_final_render — blocking a render whose generated audio … source
beat-sync-reel adjacent Generates Instagram Reels where product image cuts are synced to audio beats. Accepts audio as a local file, URL, or search query. Uses librosa for beat detection, FFmpeg Ken Burns for scene animation, and Pillow for text overlays. No AI … source
create-video-seedance-2-fal adjacent Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio via generate_audio: true, internal multi-cut … source

Source: indranilbanerjee/digital-marketing-pro

Skill Tier What it adds Upstream
c2pa-metadata adjacent Embed C2PA (Content Authenticity Initiative) provenance manifests in AI-generated marketing assets (image/video/audio/PDF). Use when: preparing AI-generated ad creative, social images, or video for EU markets to comply with EU AI Act … source

Source: infrasity-labs/dev-gtm-claude-skills

Skill Tier What it adds Upstream
blog-audio adjacent Generate audio narration of blog posts using Google Gemini TTS. Supports summary narration, full article read-aloud, and two-speaker podcast/dialogue mode with 30 voice options. Outputs MP3 with HTML5 audio embed code. Works standalone via … source

Source: jeremylongshore/claude-code-plugins-plus-skills

Skill Tier What it adds Upstream
abridge-core-workflow-a core Implement Abridge ambient clinical documentation capture-to-note pipeline. Use when building the primary encounter workflow: audio capture, real-time transcription, AI note generation, and EHR note insertion. Trigger: "abridge clinical … source
granola-common-errors core Troubleshoot common Granola errors \u2014 audio capture failures, transcription\ \ issues,\ncalendar sync problems, and integration errors. Platform-specific fixes\ \ for macOS and Windows.\nTrigger: "granola error", "granola not … source
granola-performance-tuning core Optimize Granola transcription accuracy, note quality, and processing speed. Use when improving transcription quality, reducing processing time, optimizing templates for better AI output, or tuning audio setup. Trigger: "granola … source
twinmind-performance-tuning core Optimize TwinMind transcription accuracy and speed with Ear-3 model configuration, audio quality tuning, and caching strategies. Use when implementing performance tuning, or managing TwinMind meeting AI operations. Trigger with phrases … source
granola-install-auth core Install and configure Granola AI meeting notes with calendar and audio permissions. Use when setting up Granola for the first time, connecting Google/Outlook calendars, granting macOS Screen Recording permission, or configuring Windows … source
elevenlabs-core-workflow-b adjacent Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text. Use when converting voice to another voice, generating sound effects from text, removing background noise, or transcribing audio. Trigger: … source
twinmind-security-basics adjacent Security best practices for TwinMind: on-device audio processing, encrypted cloud backups, microphone permissions, and data privacy controls. Use when implementing security basics, or managing TwinMind meeting AI operations. Trigger with … source
abridge-common-errors adjacent Diagnose and fix common Abridge clinical AI integration errors. Use when encountering EHR connectivity failures, note generation errors, audio streaming issues, or FHIR validation problems with Abridge. Trigger: "abridge error", "abridge … source
abridge-performance-tuning adjacent Optimize Abridge clinical AI integration performance for high-volume deployments. Use when reducing note generation latency, optimizing audio streaming throughput, improving FHIR push performance, or scaling for multi-site health systems. … source
assemblyai-core-workflow-a adjacent Execute AssemblyAI primary workflow: async transcription with audio intelligence. Use when transcribing audio/video files, enabling speaker diarization, sentiment analysis, entity detection, PII redaction, or content moderation. Trigger … source
assemblyai-core-workflow-b adjacent Execute AssemblyAI streaming transcription and LeMUR workflows. Use when implementing real-time speech-to-text, live captions, voice agents, or LLM-powered audio analysis with LeMUR. Trigger with phrases like "assemblyai streaming", … source
deepgram-core-workflow-a adjacent Implement production pre-recorded speech-to-text with Deepgram. Use when building audio transcription, batch processing, or implementing diarization and intelligence features. Trigger: "deepgram transcription", "speech to text", … source

Source: K-Dense-AI/claude-scientific-skills

Skill Tier What it adds Upstream
open-notebook adjacent Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. Use when organizing research materials into notebooks, ingesting diverse content sources (PDFs, videos, audio, web pages, Office … source

Source: K-Dense-AI/scientific-agent-skills

Skill Tier What it adds Upstream
open-notebook adjacent Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. Use when organizing research materials into notebooks, ingesting diverse content sources (PDFs, videos, audio, web pages, Office … source

Source: ljagiello/ctf-skills

Skill Tier What it adds Upstream
ctf-misc adjacent Provides miscellaneous CTF challenge techniques for problems that do not cleanly fit the main categories. Use for encoding puzzles, pyjails, bash jails, RF/SDR, DNS oddities, unicode tricks, esoteric languages, QR or audio puzzles, … source

Source: malob/nix-config

Skill Tier What it adds Upstream
cancel adjacent Stop any currently playing text-to-speech audio source

Source: microsoft/skills

Skill Tier What it adds Upstream
azure-ai-contentunderstanding-py adjacent Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document … source
azure-ai-openai-dotnet adjacent Azure OpenAI SDK for .NET. Client library for Azure OpenAI and OpenAI services. Use for chat completions, embeddings, image generation, audio transcription, and assistants. Triggers: "Azure OpenAI", "AzureOpenAIClient", "ChatClient", "chat … source
azure-ai-voicelive-java adjacent Azure AI VoiceLive SDK for Java. Real-time bidirectional voice conversations with AI assistants using WebSocket. Triggers: "VoiceLiveClient java", "voice assistant java", "real-time voice java", "audio streaming java", "voice activity … source
azure-ai-voicelive-py adjacent Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, … source
azure-speech-to-text-rest-py adjacent Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", … source

Source: mohitagw15856/pm-claude-skills

Skill Tier What it adds Upstream
style-fingerprint adjacent Study 3-5 documents the user actually shipped and distil a compact style card — so every skill writes in their voice, not the model's. Use when asked to learn my writing style, make outputs sound like me, build a voice profile, or when a … source
youtube-script-writer adjacent Write engaging, high-retention YouTube video scripts with visual and audio cues. Use when asked to write a YouTube script, design a video outline, draft a video hook, or structure a video narrative. Produces a polished script with multiple … source

Source: mukul975/Anthropic-Cybersecurity-Skills

Skill Tier What it adds Upstream
performing-steganography-detection adjacent Detect and extract hidden data embedded in images, audio, and other media files using steganalysis tools to uncover covert communication channels. source
detecting-deepfake-audio-in-vishing-attacks adjacent Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features (MFCC, spectral centroid, spectral contrast, zero-crossing rate) and classifying samples with machine learning models. Supports … source

Source: nexu-io/open-design

Skill Tier What it adds Upstream
audio-jingle adjacent Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3, and SFX to ElevenLabs SFX or AudioCraft. Output is one MP3/WAV file … source
od-media-generation adjacent Default reference pipeline for image, video, and audio projects — routes through media-image / media-video / media-audio atoms based on the project kind, wraps the output in a live artifact, and devloops on critique-theater until the score … source
fal-lip-sync adjacent Create talking head videos and lip sync audio to video via fal.ai. Useful for explainer avatars, multilingual dubbing previews, and social cuts. source
fal-video-edit adjacent Edit existing videos using AI — remix style, upscale, remove background, and add audio via fal.ai's hosted video models. source
hyperframes adjacent Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. Use when asked to build any HTML-based video content, add captions or subtitles synced … source

Source: NoizAI/skills

Skill Tier What it adds Upstream
sound-fx adjacent Use this skill whenever the user wants to generate sound effects, ambient audio, or short audio clips from a text description. Triggers include: any mention of 'sound effect', 'sfx', 'generate sound', 'make a sound', 'audio effect', … source
characteristic-voice adjacent Use this skill whenever the user wants speech to sound more human, companion-like, or emotionally expressive. Triggers include: any mention of 'say like', 'talk like', 'speak like', 'companion voice', 'comfort me', 'cheer me up', 'sound … source
daily-news-caster adjacent Fetches the latest news using news-aggregator-skill, formats it into a podcast script in Markdown format, and uses the tts skill to generate a podcast audio file. Use when the user asks to get the latest news and read it out as a podcast. source
chat-with-anyone adjacent Chat with any real person or fictional character in their own voice by automatically finding their speech online, extracting a clean reference sample, and generating audio replies. Also supports generating a matching voice from an uploaded … source

Source: Orchestra-Research/AI-research-SKILLs

Skill Tier What it adds Upstream
nemo-curator core GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source

Source: Orkas-AI/Orkas-VideoStudio

Skill Tier What it adds Upstream
video-craft adjacent The craft standard that makes a video GOOD, not just rendered — hooks, story pacing, visual hierarchy, type & safe zones, motion/easing, captions, audio mix, platform conventions, shot language, generation-prompt writing, and a pre-publish … source

Source: SamurAIGPT/Generative-Media-Skills

Skill Tier What it adds Upstream
muapi-media-generation adjacent Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5 source
muapi-ugc-video-factory adjacent Turn a person photo + a product photo + an optional script into a vertical 9:16 UGC-style video ad. Generates a lifestyle hero image (Nano-Banana Pro Edit), then animates it with native audio using Seedance 2.0 VIP image-to-video. source

Source: sanjay3290/ai-skills

Skill Tier What it adds Upstream
google-tts core Convert documents and text to audio using Google Cloud Text-to-Speech. Use this skill when the user wants to: narrate a document, read aloud text, generate audio from a file, convert text to speech, create a recording of documentation or … source
elevenlabs adjacent Convert documents and text to audio using ElevenLabs text-to-speech. Use this skill when the user wants to create a podcast, narrate a document, read aloud text, generate audio from a file, or convert text to speech. source

Source: silverstein/minutes

Skill Tier What it adds Upstream
minutes-setup adjacent Guided first-time setup for Minutes — download whisper model, create directories, configure audio input. Use when the user says "set up minutes", "install minutes", "first time setup", "configure minutes", "get started with minutes", "how … source

Source: transloadit/skills

Skill Tier What it adds Upstream
transform-transcribe-audio-with-transloadit adjacent One-off transcription of local audio or video files to text or subtitle files using Transloadit via the official @transloadit/node CLI. Use when the user wants speech in local media converted to .txt, .json, .srt, or .webvtt; … source

Source: Vincentwei1021/video-shotcraft

Skill Tier What it adds Upstream
video-shotcraft adjacent Create cinematic product videos from shot recipe cards, a validated template, and code/audio assets (Remotion + real page screenshots + 2.5D camera moves + beat-synced cuts + sound design). Use when the user asks to turn a frontend project … source

Source: zechenzhangAGI/AI-research-SKILLs

Skill Tier What it adds Upstream
nemo-curator core GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with … source
audiocraft-audio-generation adjacent PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation. source

Source: google/skills

Skill Tier What it adds Upstream
ima-sdk-basics adjacent Use this skill for Interactive Media Ads (IMA) SDK client-side ad insertion when you are requesting video ads client-side into websites, apps, TVs or other platforms with VAST or VMAP. Do not use for Dynamic Ad Insertion (DAI), SSAI, or … source

Source: veniceai/skills

Skill Tier What it adds Upstream
venice-chat core Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning … source
venice-audio-music adjacent Async music / audio-track generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental, voice selection, duration, language, speed, model capability probing, and … source
venice-audio-speech adjacent Generate speech from text via POST /audio/speech. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash), voices per family, output formats (mp3/opus/aac/flac/wav/pcm), streaming, … source
venice-audio-transcription adjacent Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), supported formats (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text), timestamps, and … source
venice-video adjacent Generate and transcribe videos via Venice. Covers the async /video/quote + /video/queue + /video/retrieve + /video/complete loop, text-to-video, image-to-video, video-to-video (upscale), audio input, reference images, scene and element … source