media-use
heygen-com/hyperframes · 289k installs · 59.5K stars
Skill Agent Media OS for a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs privacy, or needs a media reveal. When the host app provides its own music or sound-effect tools, use those for music and sound effects; `resolve --type bgm|sfx` needs the heygen CLI. When `HEYGEN_API_BASE` is set, HeyGen calls go through that host with no CLI sign-in.
Install
npx -y skills add heygen-com/hyperframes --skill media-use --agent claude-codeRun in a terminal. Claude loads the skill when a task calls for it. Check the repo's README before you run it.
Files
Plugin installs: Before setup or freshness commands, follow plugin execution rules when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below.
media-use
The media OS for HyperFrames: resolve · generate · operate · remember — every media type, one skill, zero context noise.
Only when HEYGEN_API_BASE is set in your environment (a host app set it and pays for HeyGen with its own key): HeyGen media is already paid for. Do not ask the person to install or sign in to the heygen CLI and do not offer its OAuth allowance; catalog search, TTS and avatar calls here go through the host. When that same host also gives you its own HeyGen tools, use those first. When a call through the host is refused, tell the person the host's message as written (it names the fix, such as adding or replacing the key in the app's Settings) and stop; do not switch to another provider unless they ask.
First run otherwise (no HEYGEN_API_BASE), when you will use HeyGen media (catalog search, TTS, avatar video): install and sign in to the heygen CLI (the free-usage path), then verify with npx hyperframes media-use resolve --doctor. Setup and providers: references/setup-providers.md.
Music and sound effects inside a host app: when the app you run in gives you its own music or sound-effect tools, use those. Without HEYGEN_API_BASE, resolve --type bgm and --type sfx search the HeyGen catalog through the heygen CLI; without it they fail and say so (sfx still answers from its bundled library).
Without HEYGEN_API_BASE, before generating a voiceover or an avatar video, tell the person: signing in to the heygen CLI with OAuth (heygen auth login --oauth) gives a free allowance for TTS voiceover and avatar videos, while an API key bills API credits.
Resolve — the one verb
npx hyperframes media-use resolve --type <type> --intent "<description>" --project <dir>
Returns one line: resolved <id> → <path> (<type>, <metadata>). All search noise stays on disk.
| Type | One-line intent |
|---|---|
bgm |
background music (HeyGen catalog via the heygen CLI, 10k+ tracks) |
sfx |
sound effects (bundled 19-file library + catalog via the heygen CLI) |
image |
photos, backgrounds (HeyGen asset search, 75k+ vectors) |
icon |
icons, symbols (transparent) |
logo |
official brand marks (theSVG → GitHub avatar → favicon; never redrawn) |
voice |
TTS voiceover (HeyGen free-usage path; optional local Kokoro) |
grade |
measured correction candidate; broad polish/stylization follows Media Treatments |
lut |
user-provided or explicitly chosen reusable validated .cube file |
Before resolving fresh, list reusable candidates with --candidates and judge fit yourself — reuse rules, all flags, ingest (--from), and adopt are in references/resolve.md.
Treat broad visual feedback as media intent
When a user explicitly asks to fix, polish, stylize, obscure, emphasize, or
reveal photographic media, read references/media-treatments.md even if they
do not name color grading or an effect. Inspect the real <img>/<video>,
choose one primary intent, then use deterministic persistence and verification.
Use a matching recipe as an optional tested seed, or inspect
hyperframes media-treatment --capabilities --json, then request one relevant
family/effect with --capability <id> and assemble a custom treatment from
canonical controls. Never load --all for ordinary authoring. A treatment may
compose correction, a preset, finishing, compatible shader effects, supported
keyframes, and optional Registry overlays. Add only source-justified bounded
tuning and compatible parts, never effects merely to make the result look more
sophisticated. Persist the final combined payload with
hyperframes media-treatment.
Use one progressively escalating workflow. For video, inspect one labeled early/middle/late contact sheet rather than reading frames separately. Apply one candidate and inspect one after-sheet for ordinary correction or polish. Escalate to individual frames or moving draft evidence only when the result is ambiguous, temporal, stylized, LUT-based, HDR/LOG-sensitive, private, or brand-critical.
For ordinary correction or polish, persist the final treatment's
preset/adjustment JSON.
Do not generate a .cube LUT merely to encode exposure, shadows, contrast, or
warmth. Use a LUT only when the user supplies one or the selected treatment
explicitly owns one. resolve --type grade --for ... --analyze is measurement
evidence, not permission to replace the chosen treatment with a generated LUT.
Do not recreate supported vignette, grain, blur, pixelate, color, or treatment
effects with CSS/SVG overlays; that bypasses Studio controls and the canonical
preview/render shader path.
Be proactive — run a media opportunity pass
The human usually can't tell which media would lift the piece. You can. When you build or review a composition, do one grounded scan and then ask once — don't silently add, and don't nag per asset.
Surface an opportunity only when a concrete signal is present:
| Signal detected | Offer |
|---|---|
| On-screen text / a script with no voiceover | TTS voiceover (audio engine) |
Emoji or a <div> styled as an icon |
resolve real icons |
| Image that is a placeholder, tiny, or upscaled-looking | a better image (and/or upscale — see references/operations.md) |
| Hard scene cuts / transitions with no sound | transition sfx |
| A piece over ~10s with no music bed | bgm |
| Footage that reads under/over-exposed or color-cast | a corrective grade (inspect it with hyperframes media-treatment --selector '#hero' --analyze --json) |
| Photographic media that feels visually flat or off-topic | one specific source-appropriate preset or custom treatment, with the intended target named |
| A meaningful media entrance/reveal that feels static | one supported seek-safe treatment animation; preserve color unless the request also justifies a preset |
Rules that keep this a help, not nagware: grounded, not generic (no signal → no suggestion); opinionated + concrete (propose the specific fix with defaults chosen — the human approves all / some / none); once per project (one consolidated ask; respect "leave it"); surface, never silently mutate (color grades especially: propose and preview — a gray-world "correction" ruins an intentional sunset or neon look).
Where to look — read only the file your task needs
| Task | Read |
|---|---|
| resolve / reuse / adopt / ingest, flags, cascade, inventory | references/resolve.md |
color grading, LUTs, smart grade (--for), grade-compare |
references/grading.md |
| voiceover / TTS, music, SFX, captions, transcription (audio engine) | references/audio.md |
| cut / reframe / transform existing media, exact error diffusion, HEVC | references/operations.md |
| source-aware creative treatments, realtime effects, overlays, reveals | references/media-treatments.md |
install + auth, provider table, RAM ladders, --local-only, --provider |
references/setup-providers.md |
| remembered preferences + frozen recipes (user memory) | references/memory.md |
| ownership matrix, usage stats, telemetry, privacy (maintainer-facing) | references/meta.md |
Facts
- Kind
- Skill
- Repo
- heygen-com/hyperframes
- Group
- Uncategorized
- Installs
- 289k
- Stars
- 59.5K
- 1find-skillsvercel-labs/skillsHelps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.3M
- 2ai-video-generationinference-sh/skillsGenerate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse1.7M
- 3ai-image-generationinference-sh/skillsGenerate AI images with GPT-Image-2.5, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2.5 Flare, GPT-Image-2.5 Sunburst, GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to image, stable diffusion, generate image, ai art, midjourney alternative, dall-e alternative, text2img, t2i, image generator, ai picture, create image with ai, generative ai, ai illustration, grok image, gemini image, gpt image, openai image, chatgpt image1.7M
- 4ai-avatar-videoinference-sh/skillsCreate AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc1.7M
- 5twitter-automationinference-sh/skillsAutomate Twitter/X with posting, engagement, and user management via inference.sh CLI. Apps: x/post-create (text and media), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Capabilities: post tweets, schedule content, like posts, retweet, send DMs, follow users, get profiles. Use for: social media automation, content scheduling, engagement bots, audience growth, X API. Triggers: twitter api, x api, tweet automation, post to twitter, twitter bot, social media automation, x automation, tweet scheduler, twitter integration, post tweet, twitter post, x post, send tweet1.7M
- 6remotion-renderinference-sh/skillsRender videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video1.6M