media-bunny
Browser-based video processing with mediabunny — read, write, convert MP4/WebM/MKV using WebCodecs API. Server integration for storage and metadata tracking.
What this skill does
# Media Bunny — Browser Video Processing for Next.js
Browser-based video processing using the mediabunny npm package and the WebCodecs API. Supports reading video metadata, converting between MP4/WebM/MKV formats, and server-side integration for storing processed videos and tracking metadata via Drizzle.
## Prerequisites
- Next.js app with `src/` directory and App Router
- Storage skill installed (for file upload/download)
- DB skill installed (for Drizzle schema and database access)
- A modern browser with WebCodecs support (Chrome 94+, Edge 94+, Opera 80+)
## Installation
```bash
bun add mediabunny
```
The `storage` and `db` dependencies are installed by their own skills.
## What Gets Created
```
src/
├── lib/
│ ├── media-bunny/
│ │ ├── types.ts # Shared types (VideoMetadata, ProcessingStatus, etc.)
│ │ └── use-media-bunny.ts # React hook for mediabunny with lazy loading
│ └── db/
│ └── schema/
│ └── videos.ts # Drizzle schema for video metadata table
├── components/
│ └── video/
│ ├── video-processor.tsx # Main video processing component (use client)
│ └── video-processor-wrapper.tsx # Dynamic import wrapper for SSR safety
└── app/
├── api/
│ └── videos/
│ ├── route.ts # GET list videos, POST create video record
│ └── [id]/
│ └── route.ts # GET single video, PATCH update, DELETE
└── video/
└── page.tsx # Test page with minimal UI
```
## mediabunny API Reference
### Core Classes
- **`Input`** — constructor: `new Input({ formats, source })` where `formats` comes from `ALL_FORMATS` and `source` is a Source instance
- Methods: `getFormat()`, `computeDuration()`, `getFirstTimestamp()`, `getTracks()`, `getVideoTracks()`, `getAudioTracks()`, `getPrimaryVideoTrack()`, `getPrimaryAudioTrack()`, `getMimeType()`, `getMetadataTags()`, `dispose()`
- **`Output`** — constructor: `new Output({ format, target })`
- Methods: `addVideoTrack(source, metadata?)`, `addAudioTrack(source, metadata?)`, `start()`, `finalize()`, `cancel()`
- Properties: `format`, `state` (`'pending'|'started'|'canceled'|'finalizing'|'finalized'`), `target`
- **`Conversion`** — static init: `Conversion.init({ input, output, video?, audio? })`
- Methods: `execute()`, `cancel()`
- Properties: `input`, `output`, `onProgress`
### Sources (reading)
`BlobSource(blob)`, `BufferSource`, `UrlSource`, `StreamSource`, `ReadableStreamSource`, `FilePathSource`
### Targets (writing)
`BufferTarget()`, `StreamTarget(stream)`, `FilePathTarget`, `NullTarget`
### Sinks
`VideoSampleSink(track)` — `getSample(timestamp)`, `samples(start, end)`, `samplesAtTimestamps(timestamps)`
### Sources (writing)
`VideoSampleSource(encodingConfig)` — `add(sample, encodeOptions?)`
### Output Formats
`Mp4OutputFormat`, `WebMOutputFormat`, `MkvOutputFormat`, `WavOutputFormat`
### Video Encoding Config
```typescript
{
codec: "avc" | "hevc" | "av1" | "vp8" | "vp9";
bitrate: number;
keyFrameInterval?: number;
sizeChangeBehavior?: string;
alpha?: boolean;
bitrateMode?: string;
latencyMode?: string;
}
```
### Codec-Container Compatibility
| Container | Video Codecs | Audio Codecs |
|-----------|-------------|--------------|
| MP4 | avc, hevc, av1 | aac, mp3, flac |
| WebM | vp8, vp9, av1 | opus, vorbis |
| MKV | avc, hevc, av1, vp8, vp9 | aac, opus, mp3, vorbis, flac |
## Setup Steps
### Step 1: Create `src/lib/media-bunny/types.ts`
```typescript
export type ProcessingStatus = "pending" | "processing" | "completed" | "failed";
export type VideoMetadata = {
id: string;
fileName: string;
format: string;
duration: number;
width: number;
height: number;
codec: string;
fileSize: number;
storageKey: string | null;
status: ProcessingStatus;
createdAt: Date;
updatedAt: Date;
};
export type VideoProcessingOptions = {
outputFormat: "mp4" | "webm";
videoCodec?: import("mediabunny").VideoCodec;
audioCodec?: import("mediabunny").AudioCodec;
videoBitrate?: number;
audioBitrate?: number;
};
export type VideoInfo = {
duration: number;
width: number;
height: number;
format: string;
codec: string;
};
export type UseMediaBunnyReturn = {
isLoading: boolean;
isSupported: boolean;
error: string | null;
getVideoInfo: (file: File) => Promise<VideoInfo>;
convertVideo: (
file: File,
options: VideoProcessingOptions,
onProgress?: (progress: number) => void
) => Promise<Blob>;
};
```
### Step 2: Create `src/lib/media-bunny/use-media-bunny.ts`
```typescript
"use client";
import { useCallback, useEffect, useRef, useState } from "react";
import type {
UseMediaBunnyReturn,
VideoInfo,
VideoProcessingOptions,
} from "./types";
type MediaBunnyModule = typeof import("mediabunny");
export function useMediaBunny(): UseMediaBunnyReturn {
const [isLoading, setIsLoading] = useState(true);
const [isSupported, setIsSupported] = useState(false);
const [error, setError] = useState<string | null>(null);
const moduleRef = useRef<MediaBunnyModule | null>(null);
useEffect(() => {
const supported =
typeof VideoDecoder !== "undefined" &&
typeof VideoEncoder !== "undefined";
setIsSupported(supported);
if (!supported) {
setIsLoading(false);
setError("WebCodecs API is not supported in this browser.");
return;
}
import("mediabunny")
.then((mod) => {
moduleRef.current = mod;
setIsLoading(false);
})
.catch((err: unknown) => {
const message =
err instanceof Error ? err.message : "Failed to load mediabunny";
setError(message);
setIsLoading(false);
});
}, []);
const getVideoInfo = useCallback(async (file: File): Promise<VideoInfo> => {
const mb = moduleRef.current;
if (!mb) {
throw new Error("mediabunny is not loaded yet");
}
const source = new mb.BlobSource(file);
const input = new mb.Input({
formats: mb.ALL_FORMATS,
source,
});
try {
const format = await input.getFormat();
const duration = await input.computeDuration();
const videoTrack = await input.getPrimaryVideoTrack();
const width = videoTrack?.displayWidth ?? 0;
const height = videoTrack?.displayHeight ?? 0;
const codec = videoTrack?.codec ?? "unknown";
return {
duration,
width,
height,
format: format?.name ?? "unknown",
codec,
};
} finally {
input.dispose();
}
}, []);
const convertVideo = useCallback(
async (
file: File,
options: VideoProcessingOptions,
onProgress?: (progress: number) => void
): Promise<Blob> => {
const mb = moduleRef.current;
if (!mb) {
throw new Error("mediabunny is not loaded yet");
}
const source = new mb.BlobSource(file);
const input = new mb.Input({
formats: mb.ALL_FORMATS,
source,
});
const target = new mb.BufferTarget();
const format =
options.outputFormat === "mp4"
? new mb.Mp4OutputFormat()
: new mb.WebMOutputFormat();
const output = new mb.Output({ format, target });
const videoCodec =
options.videoCodec ?? (options.outputFormat === "mp4" ? "avc" : "vp9");
const audioCodec =
options.audioCodec ?? (options.outputFormat === "mp4" ? "aac" : "opus");
try {
const conversion = await mb.Conversion.init({
input,
output,
video: {
codec: videoCodec,
bitrate: options.videoBitrate ?? 2_000_000,
},
audio: {
codec: audioCodec,
bitrate: options.audioBitrate ?? 128_000,
},
});
if (onProgress) {
conversion.onProgress = onProgress;
}
await conversion.execute();
const mimeType =
Related in Image & Video
watch
IncludedWatch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or Whisper API fallback), and hands the result to Claude so it can answer questions about what's in the video.
physical-ai-defect-image-generation
IncludedUse when the user wants to orchestrate defect image generation, run associated setup, or handle outputs on OSMO. The Day 0 path handles cold-start with USD-to-ROI, image-edit augmentation, and AnomalyGen to create initial PCBA datasets. The Day 1 path performs inference and labeling on real images. This skill helps with first-time asset setup, creation of finetuning checkpoints, and configuring deployment. Trigger keywords: defect image generation, dig workflow, dig pipeline, defect image detection workflow, aoi pipeline, aoi anomalygen, usd2roi anomalygen, day 0 pcba, day 1 pcba, day 1 real-photo alignment, day 1 manual roi, metal surface anomaly, glass defect, anomalygen finetune, setup_pcb, setup_metal, setup_glass, setup_pretrained, dig setup, dig datasets, dig pretrained checkpoint, dig image-edit endpoint.
accelint-react-best-practices
IncludedReact performance optimization and best practices. ALWAYS use this skill when working with any React code - writing components, hooks, JSX; refactoring; optimizing re-renders, memoization, state management; reviewing for performance; fixing hydration mismatches; debugging infinite re-renders, stale closures, input focus loss, animations restarting; preventing remounting; implementing transitions, lazy initialization, effect dependencies. Even simple React tasks benefit from these patterns. Covers React 19+ (useEffectEvent, Activity, ref props). Triggers - useEffect, useState, useMemo, useCallback, memo, inline components, nested components, components inside components, re-render, performance, hydration, SSR, Next.js, useDeferredValue, combined hooks.
elevenlabs-agents
IncludedBuild conversational AI voice agents with ElevenLabs Platform using React, JavaScript, React Native, or Swift SDKs. Configure agents, tools (client/server/MCP), RAG knowledge bases, multi-voice, and Scribe real-time STT. Use when: building voice chat interfaces, implementing AI phone agents with Twilio, configuring agent workflows or tools, adding RAG knowledge bases, testing with CLI "agents as code", or troubleshooting deprecated @11labs packages, Android audio cutoff, CSP violations, dynamic variables, or WebRTC config. Keywords: ElevenLabs Agents, ElevenLabs voice agents, AI voice agents, conversational AI, @elevenlabs/react, @elevenlabs/client, @elevenlabs/react-native, @elevenlabs/elevenlabs-js, @elevenlabs/agents-cli, elevenlabs SDK, voice AI, TTS, text-to-speech, ASR, speech recognition, turn-taking model, WebRTC voice, WebSocket voice, ElevenLabs conversation, agent system prompt, agent tools, agent knowledge base, RAG voice agents, multi-voice agents, pronunciation dictionary, voice speed control, elevenlabs scribe, @11labs deprecated, Android audio cutoff, CSP violation elevenlabs, dynamic variables elevenlabs, case-sensitive tool names, webhook authentication
humanizer
IncludedHumanize AI-generated text by detecting and removing patterns typical of LLM output. Rewrites text to sound natural, specific, and human. Uses 28 pattern detectors, 560+ AI vocabulary terms across 3 tiers, and statistical analysis (burstiness, type-token ratio, readability) for comprehensive detection. Use when asked to humanize text, de-AI writing, make content sound more natural/human, review writing for AI patterns, score text for AI detection, or improve AI-generated drafts. Covers content, language, style, communication, and filler categories.
generating-mermaid-diagrams
IncludedSalesforce architecture diagrams using Mermaid with ASCII fallback. Use this skill when generating text-based diagrams for Salesforce architecture, OAuth flows, ERDs, integration sequences, or Agentforce structure. TRIGGER when: user says "diagram", "visualize", "ERD", or asks for sequence diagrams, flowcharts, class diagrams, or architecture visualizations in Mermaid. DO NOT TRIGGER when: user wants PNG/SVG image output (use generating-visual-diagrams), or asks about non-Salesforce systems.