Photoshop-, After Effects-, and Blender-class video tools as one npm package that AI agents drive with code.
⚠️ Beta: gitframes is under active development. APIs may change between releases and some features may be incomplete or unstable.
Gitframes is built for coding agents. It packs the work people usually split across three desktop apps (Photoshop-grade compositing and VFX, After Effects-style motion, typography and keyframing, and Blender-style 3D scenes, cameras and models) into one lightweight npm package. Your agent writes a TypeScript composition, checks frames, and renders an MP4, and nobody has to install or license a multi-gigabyte creative suite.
Code-first video as pure software engineering — no headless browser, no DOM reflow, no screenshot pipeline. Renders directly on GPU hardware via Dawn / WebGPU / Metal / Vulkan in Node.js and modern WebGPU browsers.
Every frame of these films is rendered by gitframes from TypeScript in examples/. Click a still to watch it on YouTube.
[!NOTE] Using an AI coding agent? Install the gitframes skills in one line.
Claude Code
/plugin install gitframesAny other agent (Codex, Cursor, Hermes, Gemini CLI, Copilot, and more)
npx skills add gatewai-dev/gitframesSee Agent Skills & Plugins for details.
Modern automated video generation is usually constrained by the architectures of general-purpose web browsers: process overhead, non-deterministic DOM layout reflows, and slow screenshot capture. Gitframes treats video composition as software engineering:
| Pillar | What it means | |
|---|---|---|
| 🚀 | Zero Headless-Browser Overhead | No Puppeteer, no Chromium IPC, no page.screenshot(). Gitframes talks straight to native GPU devices via Dawn/WebGPU and hardware-encodes with @napi-rs/webcodecs. |
| 🎯 | Deterministic Frame-Accurate Clock | Absolute frame clocks, discrete sample points, and frame-accurate audio BeatGrids. No floating timers, no drift, no dropped frames. |
| 🔠 | Analytic, Resolution-Independent Type | The Slug algorithm evaluates glyph contours per-pixel in WGSL — no texture atlases, no scaling artifacts, razor-sharp from 10 px to 10,000 px. |
| 🎨 | Photoshop-Grade Tonal & Spatial VFX | 50+ modular GPU shaders: Curves, Levels, Selective Color, 3D LUTs, Halftone, Film Grain, Unsharp Mask, Mesh Warp, and Screen-Space Relighting. |
| 🧊 | Unified 3D & 2D Depth Compositing | Nest 2D flex/box trees inside 3D homography planes, multiplane rigs, and meshes (OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF), with PBR glass and SSAO. |
| 🔊 | Built-in Procedural Audio DSP | Multi-track soundtracks, deterministic procedural transition SFX (whoosh, impact, riser), and reactive signals that drive visuals from audio. |
| 👁️ | On-Device Neural Vision | Object tracking, instance segmentation, multi-person pose, and person mattes from Apache-2.0 ONNX models — feeding reactive signals without a round trip to disk. |
| ☁️ | Cloud-Native & CI/CD Ready | ~200–400 MB RAM per worker (vs. 2–4 GB for Chromium), ideal for serverless GPU render clusters (AWS G4/G5, Modal, RunPod, Kubernetes). |
Developers generating video programmatically commonly weigh Remotion (React/Chromium) or Hyperframes (Canvas2D/SVG web animation). The matrix below compares the fundamental engineering dimensions.
| Capability / Dimension | Gitframes | Remotion | Hyperframes |
|---|---|---|---|
| Underlying Engine | Native WebGPU (WGSL compute & render pipelines via Dawn / Metal / Vulkan) | Chromium / Puppeteer (React DOM, HTML/CSS layout) | Canvas2D / WebGL / SVG (browser or Node Skia) |
| Rendering Architecture | Direct hardware framebuffer rendering & hardware video encoding (@napi-rs/webcodecs) | Spawns headless Chrome; captures frames via CDP / page.screenshot() | Software or hardware 2D canvas context |
| Throughput | 60–120+ FPS (real-time to faster-than-real-time GPU execution) | 5–20 FPS (DOM reflow, IPC, rasterization) | 20–40 FPS (CPU draw commands / JS) |
| Memory Footprint | ~200–400 MB per render (zero browser) | 1.5–4.0 GB+ per worker (Chromium + V8 DOM heap) | ~500 MB–1 GB (Skia/Canvas bindings) |
| Typography Engine | Slug GPU — analytic Bézier evaluation in WGSL, infinite zoom, After Effects selectors | Browser DOM text (CSS fonts, rasterized, blurry under 3D transforms) | Canvas2D / path text (CPU-rasterized glyphs) |
| 2D VFX & Post-Processing | 50+ WebGPU shaders (Curves, Levels, Selective Color, 3D LUT, Film Grain, Halftone, Liquify, PBR Glass, Relight) | CSS Filters or custom WebGL canvas wrappers | Basic Canvas2D composites and 2D filters |
| 3D Graphics & Depth | Native 3D scene graph — LookAt/Turntable camera, multiplane, skinning (OBJ/FBX/glTF), SSAO, PCSS, DoF | None built-in (embed Three.js/Fiber inside React DOM) | Minimal 2.5D layers; no unified mesh pipeline |
| Motion Blur & Physics | Physical 180° shutter velocity buffers in MRT + closed-form spring kinematics | CSS transitions / JS interpolation; synthetic blur hacks | Frame interpolation or manual multipass |
| Audio Engine & DSP | Native audio DSP & procedural SFX (multi-track mixing, beat grids, reactive signals) | <Audio> playback; basic volume curves | Basic static audio playback |
| Charts & Data Viz | Layer.chart — line, area, bar, scatter, candlestick, pie and donut charts built from native vector nodes, with staggered reveal animations | DOM chart libraries (Recharts, Chart.js) | Custom canvas draw operations |
| AI & Computer Vision | On-device ONNX vision — COCO-80 detection + instance masks (RTMDet-Ins), COCO-17 pose (RTMO), person mattes (Selfie Segmenter); WebGPU tensor conditioning (Canny, depth-to-normals, optical flow, deflicker) | External pre-rendered assets; no native GPU tensor conditioning | External pre-rendered assets |
| Headless Verification | FrameGrid contact sheets, single-frame snapshots, Skia MSE pixel-invariant assertions | Playwright/Puppeteer visual snapshots | Manual frame inspection / canvas diffing |
| Docker / Cloud Portability | Compact (~500 MB slim image with native GPU/Vulkan drivers) | Heavy (~2–3 GB with Chromium, fonts, X11/Mesa) | Moderate container size |
Traditional text relies on CPU rasterization or low-res SDF atlases that soften under 3D camera sweeps. Gitframes integrates the Slug algorithm (SlugPipeline):
square, ramp_up, ramp_down, triangle, smooth), easeHigh/easeLow curves, and seeded PRNG character shuffling (TextAnimator).TypewriterAnimator).evaluateVolumetricFormation).A comprehensive suite of professional image/video shader nodes in nodes/ and packages/webgpu-renderers:
ApplyLUT).PBRGlass).Camera3D) calibrated so z = 0 matches 2D canvas pixel coordinates 1:1.Layer3D.cube, carousel, prism, plane, grid with unified depth-buffer testing.rg16float MRT buffers..audio media nodes with frame-exact lifecycle control.renderSfx, mixSfxInto, softLimit).mixAudioTracks and encodeStereoWav.Signal.builder) or audio analysis.Layer.chart builds line, area, bar (grouped or stacked), scatter, candlestick, pie and donut charts. d3 computes the scales, ticks and geometry; every bar, line, slice and label is an ordinary box, path or text node:
animate: { start, duration, stagger, ease }, or animate: false).Layer.chart(
{
type: "bar",
width: 900,
height: 480,
categories: ["Q1", "Q2", "Q3", "Q4"],
series: [
{ name: "Revenue", data: [12, 19, 24, 31] },
{ name: "Costs", data: [8, 11, 13, 15] },
],
yAxis: { format: "$,.0f" },
animate: { start: 10, duration: 30 },
},
{ position: "absolute", x: 120, y: 200 },
);
@gitframes/vision runs ONNX models via onnxruntime-node (CPU) or onnxruntime-web (WebGPU) and wires every result into the same reactive signal surface the rest of Gitframes consumes.
[!TIP] Lazy by construction.
VisionRunner.create(),comp.withVision(...)andVisionNode.attach(...)perform zero I/O — no downloads, no sessions, no file probes. A model is fetched the first time a task actually runs. To warm up ahead of time, callawait runner.preload(["detect", "pose"])(orawait vision.ready()on an attached node).
Every model is Apache-2.0, pinned to an immutable Hugging Face revision, and verified by SHA-256 after download.
| Task | Option | Model | Output |
|---|---|---|---|
| Detect | enableDetection | RTMDet-Ins t/s/m (OpenMMLab) | COCO-80 boxes + scores, tracked over time |
| Segment | enableSegmentation | RTMDet-Ins (same forward pass as detect) | Soft per-instance masks, frame-aligned |
| Pose | enablePose | RTMO t/s/m (OpenMMLab) | 17 COCO keypoints + visibility per person |
| Matte | enableMatte | MediaPipe Selfie Segmenter (Google) | Fast person-vs-background alpha for portrait / webcam framing |
variant: "t" | "s" | "m" (default "s"; ~24 / 43 / 116 MB for RTMDet-Ins). Tune confidence and a COCO classes filter per composition. On CPU, a 2K frame takes roughly 200–340 ms to detect + segment, ~120 ms for pose and ~20 ms for the matte with "s".matteSource: "instance", the default).mask / matte / crop modes merge every comparably sized instance that overlaps the main subject, so a flowing dress or a held instrument stays attached to the person, while a tunnel or window framing them does not.$GITFRAMES_MODELS_DIR (default ~/.cache/gitframes/models). Point baseUrl or GITFRAMES_MODELS_BASE_URL at your own mirror for air-gapped or CI renders.PoseSkeletonRenderer).SegmentationTexturePool).TemporalObjectTracker) assigns stable trackIds via IoU association, with configurable minHits, positionSmoothing, and velocity-based coasting for up to maxMissedFrames (default 15) so a transient miss holds the track instead of flashing.pose-track-matcher) binds keypoints to the right track by id, then by spatial IoU fallback.comp.analyzeVisionSequence(src, { tasks, categories }) decodes frames through the mediabunny pipeline, tracks them, and returns a zod-serializable report (per-track frame ranges, mean speed, sampled center paths, per-class presence/confidence, mean mask coverage, model download bytes/timing) (analyzeSequence).Every tracked entity is exposed as reactive ProgrammaticSignals that animate layers and shader uniforms:
| Group | Highlights |
|---|---|
objects | get(trackId), byCategory(cat, rank), primary, count, hasCategory, detectedCategories |
objects.*.bounds | x/y/width/height, screenX/screenY/screenWidth/screenHeight, aspectRatio, area |
objects.*.anchors | 9 anchors (corners, edges, center) ready for pinning |
objects.*.kinematics | vx, vy, speed, acceleration, headingRad/Deg |
objects.*.pose | All 17 COCO keypoints, plus hasPose, wristSpeed, handRaised, bodyTiltAngle |
masks | get(trackId), subject, count; per-mask area, coverage, solidity, bboxFill |
segmentation | subject, humanSilhouette, instanceMasks, matte.coverage, GPU stencilTexture |
classes | Per-class count, maxConfidence, present, primary, plus a detection histogram |
| Tensors | poseLandmarksTensor [17,3], objectsTensor [16,8], masksTensor [16,2], histogramTensor [80] |
Project normalized landmarks to screen space with a configurable camera FOV, then bind any node to a track or landmark (SpatialLandmarkTransformer, spatial-pin):
pinToObject(track, { anchor, offsetX/Y/Z, matchWidth, matchHeight, smoothFrames, hideWhenLost })pinToLandmark(coord, { offsetX/Y/Z })comp.addSubjectSandwich({ source, behind, feather, fit }) cuts the foreground subject out and places typography/graphics behind them.comp.addSmartFraming({ source, target, targetAspect, damping, leadHeadroom }) auto-crops 16:9 → 9:16 while tracking target.comp.addSubjectOutline(vision.segmentation.subject, { source, color, width, blur }) strokes the segmented boundary as an audio-reactive contour glow.layer.blurRegion(track, { strength }) blurs faces, plates, or any detected class.passthrough, mask, matte, crop, skeleton, boxes, tracking; pick the cutout alpha with matteSource: "instance" | "selfie", and optionally keyBackground to grow the subject into connected foreground.@gitframes/vision/schemas) so the hot path stays zod-free. Unknown or removed options are rejected, not silently ignored.vision.summary(frame) returns a deterministic, serializable snapshot (objects, classes, masks) safe to call inside a frame hook.@gitframes/vision/web re-exports the engine plus createWebGPUProvider() / hasWebGPU(); onnxruntime-web is an optional lazy peer.skia-canvas to verify shader math, font coverage, and Mean Squared Error (MSE) temporal deltas.comp.renderFrameGrid(...) outputs sequential-frame contact sheets for instant review of easing, kinetic type, and transitions.startPreview({ entry, export }) serves a localhost WebGPU player that loads the composition's own module and renders every frame live in the browser. Nothing is streamed: the server only hands over the bundle, the project's assets, and the soundtrack mixed by the export engine.startPreview serves the page and returns its URL instead of opening a browser, so an agent can show it in its own pane (Claude Code, Codex); pass open: true to open the system browser.import { startPreview } from "gitframes";
const session = await startPreview(
{ entry: new URL("./film.ts", import.meta.url), export: "buildFilm" },
{ title: "gitframes film" },
);
console.log(`Preview at ${session.url}`);
await session.closed; // serves until its tab closes or a newer preview takes over
Managed with pnpm workspaces and turbo:
gitframes/
├── packages/
│ ├── gitframes/ # Unified SDK (Composition, Layer, LayerAnimation, Signal, effects)
│ ├── core/ # Core AST, Effect base class, VirtualMediaData, vision types
│ ├── compositions/ # Layout engine, Flex/Box AST compiler, timeline evaluator
│ ├── webgpu-renderers/ # WGSL shaders, Slug text engine, 3D renderer, camera, lights, materials
│ ├── tensor-webgpu/ # WebGPU compute pipelines (Canny, depth-to-normals, flow, deflicker, landmarks)
│ ├── vision/ # ONNX vision engine: detect, segment, pose, matte, tracking, signals
│ ├── renderer/ # Headless Node.js WebGPU renderer via Dawn, WebCodecs, skia-canvas
│ ├── renderers/ # Higher-level render orchestration
│ ├── media/ # Media decoding / encoding adapters
│ ├── node-sdk/ # Node renderer contracts and result schemas
│ ├── server-utils/ # Server infrastructure, storage, asset caches
│ └── client-utils/ # Shared browser utilities
├── nodes/ # 58+ specialized domain nodes (VFX, audio, layout, node-vision)
├── apps/
│ └── renderer-service/ # Production HTTP / gRPC rendering microservice container
├── examples/ # Reference compositions and films
├── plugins/gitframes/ # Agent plugin: skills only (setup, compose, effects, render)
└── scripts/ # Build, release, and plugin validation tooling
pnpm add gitframes
Requirements: Node.js ≥ 22. Gitframes uses native GPU acceleration via Dawn / WebGPU or Vulkan.
import { Composition, Layer, LayerAnimation } from "gitframes";
// 1. Initialize a 1080p60 composition
const comp = new Composition({
width: 1920,
height: 1080,
fps: 60,
durationFrames: 180, // 3 seconds
backgroundColor: "#090a0f",
fonts: ["assets/fonts/Inter.ttf", "assets/fonts/SpaceGrotesk.ttf"],
});
// 2. Define physical snap-overshoot animations
const cardEntrance = LayerAnimation.create()
.fadeIn(0, 20, "power2.out")
.fromTo("y", 60, 0, { start: 0, end: 35, ease: "back.out(1.5)" })
.fromTo("scale", 0.92, 1.0, { start: 0, end: 35, ease: "back.out(1.2)" });
// 3. Assemble a responsive flex-layout card
const heroCard = Layer.box({
width: 720,
height: 380,
background: "#141721",
borderRadius: 24,
borderColor: "#262b3d",
borderWidth: 1.5,
padding: 32,
children: [
Layer.flex({
dir: "column",
gap: 16,
children: [
Layer.text("GITFRAMES ENGINE", {
fontSize: 16,
fontWeight: 700,
fill: "#6366f1",
letterSpacing: 2.0,
}),
Layer.text("Next-Gen WebGPU Motion", {
fontSize: 48,
fontWeight: 700,
fill: "#f8fafc",
fontFamily: "SpaceGrotesk",
}),
Layer.text("Direct hardware video composition without headless browser overhead.", {
fontSize: 20,
fill: "#94a3b8",
lineHeight: 28,
}),
],
}),
],
}).animate(cardEntrance);
comp.add(heroCard);
import { Composition, Layer, Layer3D, CameraAnimation, Light } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60, durationFrames: 300 });
// 1. LookAt 3D camera with a continuous orbit
const cameraAnim = CameraAnimation.camera().orbit({
azimuth: { from: -30, to: 30 },
elevation: { from: 15, to: 15 },
radius: { to: 1200 },
start: 0,
end: 300,
});
comp.add(
Layer.camera({ x: 960, y: 540, z: -1000, targetX: 960, targetY: 540, targetZ: 0 }).animate(cameraAnim)
);
// 2. Studio lighting
comp.add(Light.ambient("#ffffff", 0.4));
comp.add(Light.directional({ color: "#e0e7ff", intensity: 1.2, x: 500, y: -800, z: -600 }));
// 3. 3D model with skeletal animation
comp.add(
Layer.glb("assets/models/character.glb", {
x: 960,
y: 640,
z: 0,
scale: 2.5,
material: "lit",
loop: true,
})
);
// 4. 3D prism layout carousel
comp.add(
Layer3D.carousel({
radius: 400,
items: [
Layer.box({ width: 280, height: 180, background: "#1e293b", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#334155", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#0f172a", borderRadius: 16 }),
],
})
);
import { Composition, Layer, LayerAnimation, Signal, renderSfx, mixSfxInto, softLimit } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
const totalFrames = 240;
// 1. Soundtrack layer
comp.addAudio(Layer.audio("assets/score.mp3", { volume: 0.9, durationFrames: totalFrames }));
// 2. Frame-accurate procedural SFX on the beat grid
const bed: [Float32Array, Float32Array] = [
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
];
mixSfxInto(bed, [
renderSfx({ type: "whoosh", atBar: 0.79, volume: 0.5 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 1 }),
renderSfx({ type: "impact", atBar: 1.0, volume: 0.8 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 2 }),
]);
softLimit(bed);
// 3. Tempo signal (120 BPM = 2 Hz)
const beatPulse = Signal.builder({ type: "sawtooth", frequency: 2, amplitude: 0.08, offset: 1.0 });
// 4. Bind it to visuals
const reactiveCard = Layer.box({ width: 400, height: 250, background: "#1c202e", borderRadius: 20 })
.animate(
LayerAnimation.create()
.signal("scale", beatPulse, { multiplier: 1.0, offset: 0.0 })
.fromTo("opacity", 0, 1, { start: 0, end: 15, ease: "power2.out" }),
);
comp.add(reactiveCard);
import { Composition, FilmGrain, Vignette, ColorBalance } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
// Whole-composition cinematic grade + film emulsion
comp.apply(new Vignette({ strength: 0.28, radius: 0.85 }));
comp.apply(new FilmGrain({ strength: 0.06, size: 1.5, animated: true }));
comp.apply(
new ColorBalance({
shadows: { cyanRed: 0, magentaGreen: 2, yellowBlue: 6 },
highlights: { cyanRed: 4, magentaGreen: 1, yellowBlue: -2 },
}),
);
import { Composition, Layer, Vignette } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 30 });
// Run vision on the whole composition. Models download lazily on first use.
const vision = comp.withVision({
enableDetection: true,
enableSegmentation: true,
enablePose: true,
variant: "s",
confidence: 0.35,
});
// Pin a caption to the primary tracked subject (smoothing + auto-hide when lost)
comp.add(
Layer.text("SUBJECT 01", { fontSize: 40, fill: "#f8fafc" }).pinToObject(
vision.objects.primary,
{ anchor: "topCenter", offsetY: -48, smoothFrames: 5, hideWhenLost: true },
),
);
// Drive a shader uniform from a reactive signal — here, subject mask coverage
comp.add(
Layer.box({ width: 1920, height: 1080, background: "#000000" }).withEffect(
new Vignette({ strength: vision.segmentation.subject.coverage, radius: 0.9 }),
),
);
// Or use the one-liners for the common editorial moves:
// comp.addSubjectSandwich({ source: "assets/dancer.mp4", behind: [headline], feather: 4 });
// comp.addSmartFraming({ source: "assets/action.mp4", target: vision.objects.primary, targetAspect: 9 / 16 });
// comp.addSubjectOutline(vision.segmentation.subject, { source: "assets/character.mp4", color: "#FF5A1F", width: 6 });
// Inspect a source before authoring: one-shot, ffmpeg-free, zod-serializable report
const report = await comp.analyzeVisionSequence("assets/street.mp4", {
tasks: ["detect", "pose"],
categories: ["person"],
});
console.log(report.tracks.map((t) => `${t.category}#${t.trackId} ${t.frames.join("–")}`));
Standalone runner (no composition):
import { VisionRunner } from "@gitframes/vision";
const runner = VisionRunner.create({ variant: "s", confidence: 0.3 }); // zero I/O
const frame = { data: rgba, width: 1920, height: 1080 };
const boxes = await runner.detect(frame); // downloads RTMDet-Ins on first call
const { masks } = await runner.segment(frame); // same forward pass, no second inference
const { people } = await runner.pose(frame); // RTMO, COCO-17 keypoints
runner.close();
In the browser (WebGPU EP):
import { VisionRunner, createWebGPUProvider, hasWebGPU } from "@gitframes/vision/web";
if (hasWebGPU()) {
const runner = VisionRunner.create({ provider: createWebGPUProvider() });
}
import { buildMyComposition } from "./my-composition.js";
const comp = await buildMyComposition();
// 1. Single frame to a PNG buffer for visual inspection
const frameBuffer = await comp.renderFrame({ frame: 45 });
// 2. Contact-sheet grid of 12 sequential frames
const gridBuffer = await comp.renderFrameGrid({
startFrame: 0,
endFrame: 120,
stepFrames: 10,
cellWidth: 320,
showLabels: true,
});
// 3. Final hardware-encoded MP4 with mixed audio
const { filePath } = await comp.renderVideo({
outputPath: "output/final-product-film.mp4",
quality: "high",
concurrency: 4,
});
console.log(`Video rendered successfully to: ${filePath}`);
THEME for colors, type, radii, and spacing. Never hardcode magic hex values or ad-hoc margins.color * opacity * alpha) must use srcFactor: "one" in their blend state ({ srcFactor: "one", dstFactor: "one-minus-src-alpha", operation: "add" }). Never use srcFactor: "src-alpha" for premultiplied output — squaring alpha darkens fades into murky gray.back.out(1.4–1.7) for snap-overshoot entrances, spring / expo.out for decelerating motion, power2.in for exits. Reserve linear for infinite spinners and time counters.skia-canvas pixel sampling in Vitest before shipping.Gitframes ships agent skills that teach Claude, Codex, and other coding agents how to write, render, and check compositions. The plugin (gitframes) is listed in Anthropic's official plugin directory and contains only skills — no MCP servers, hooks, or commands. Every other agent gets the same skills through the skills CLI.
| Skill | Use it for |
|---|---|
gitframes | Starting a project: install from npm, scaffold a composition and render script, first verified render |
gitframes-compose | Compositions, layer trees, layout, animation and easing, beat grids, film structure |
gitframes-effects | Effect classes, the unified section architecture, premultiplied-alpha invariants, vision conditioning |
gitframes-render | Headless rendering, FrameGrid inspection, pixel probes, MP4 delivery checks |
Once installed, skills load automatically when a task matches (e.g. "add a film-grain pass to this scene" or "render a frame grid of intro.ts").
The plugin is instructions only. It bundles no executables, MCP servers, hooks, or package launchers, and it sends no data anywhere. The skills tell your agent to add the gitframes npm package to your project and how to use it. When that code uses on-device vision, the SDK downloads the pinned model weights from Hugging Face on first use (see On-Device Vision). Nothing else leaves your machine.
/plugin install gitframes
Or from your shell:
claude plugin install gitframes@claude-plugins-official
It installs from Anthropic's official marketplace, which Claude Code adds for you, so there is no marketplace step, and plugins from it update automatically. Afterwards, restart Claude Code or run /reload-plugins. /plugin commands need an interactive claude terminal; in the desktop app's Code tab, use the shell form or + > Plugins > Add plugin and pick Gitframes.
Add --scope project to the shell form to record the plugin in .claude/settings.json for the whole team.
Enable it for everyone in your repo. Commit this to .claude/settings.json; Claude Code prompts teammates to install it when they trust the folder:
{
"enabledPlugins": {
"gitframes@claude-plugins-official": true
}
}
Straight from this repository (tracks main instead of the directory release):
/plugin marketplace add gatewai-dev/gitframes
/plugin install gitframes@gitframes-plugins
The skills CLI installs the skills into any of 70+ agents, including Codex, Cursor, Hermes, Gemini CLI, GitHub Copilot, Windsurf, OpenCode, and Goose:
npx skills add gatewai-dev/gitframes
It detects the agents on your machine and asks where to install. To choose them yourself, pass -a once per agent, add -g to install for your user instead of this project, and -y to skip the prompts:
npx skills add gatewai-dev/gitframes -a codex -a cursor -a hermes-agent -g -y
Keep them current with npx skills update, and remove them with npx skills remove.
Or copy the folders by hand: put plugins/gitframes/skills/<name>/ into .claude/skills/, .agents/skills/, or ~/.agents/skills/. VS Code / Copilot / Cursor / Kiro can load the portable plugin.json through their plugin UI.
The plugin lives in plugins/gitframes/ so installs carry only the skills; users get the engine from npm. Two manifests there describe it: plugin.json (portable Agent Plugins 1.0, which also carries the OpenAI listing metadata) and .claude-plugin/plugin.json. The marketplace catalog is .claude-plugin/marketplace.json. The portable field set is closed — client-specific fields go in that client's manifest, not in plugin.json. The version in both follows the gitframes package: pnpm run version:packages syncs it after changeset version (or run pnpm run sync:plugin-version on its own), since clients use it to decide when to update.
Inside this repository, Claude Code and other agents pick up skills through the symlinks in .agents/skills/ and .claude/skills/. Skills live only under plugins/gitframes/skills/; never copy them elsewhere. pnpm run check:plugins validates manifests, skill frontmatter, marketplace catalogs, symlinks, and the generated effects catalog. pnpm run sync:effects-catalog regenerates the gitframes-effects catalog after any Effect class change.
The examples/ directory holds production-grade reference compositions:
| Example | What it demonstrates |
|---|---|
19_gitframes_film | The 30-second master brand film — full pipeline, audio, VFX, 3D |
21_full_circle | Multi-scene narrative composition |
22_gitframes_launch | Launch/product-motion composition |
Gitframes uses pnpm (10+) and turbo for orchestration.
# Install
pnpm install
# Build all packages
pnpm build
# Run conformance tests
pnpm test
# Check the vision models end to end (downloads ~380 MB of weights once)
pnpm --filter @gitframes/vision test:models
# Render a specific showcase example
pnpm --filter @gitframes/example-21-full-circle render
# Render the master brand film
cd examples/19_gitframes_film && pnpm render
An optimized Dockerfile.renderer deploys the renderer service into cloud GPU clusters:
docker build -t gitframes-renderer -f Dockerfile.renderer .
Gitframes is open-source software licensed under Apache-2.0. The vision models it downloads on demand — RTMDet-Ins and RTMO (OpenMMLab) and the Selfie Segmenter (Google) — are also Apache-2.0; see registry.ts for exact sources and checksums.
Photoshop-, After Effects-, and Blender-class video tools as one npm package that AI agents drive with code.
⚠️ Beta: gitframes is under active development. APIs may change between releases and some features may be incomplete or unstable.
Gitframes is built for coding agents. It packs the work people usually split across three desktop apps (Photoshop-grade compositing and VFX, After Effects-style motion, typography and keyframing, and Blender-style 3D scenes, cameras and models) into one lightweight npm package. Your agent writes a TypeScript composition, checks frames, and renders an MP4, and nobody has to install or license a multi-gigabyte creative suite.
Code-first video as pure software engineering — no headless browser, no DOM reflow, no screenshot pipeline. Renders directly on GPU hardware via Dawn / WebGPU / Metal / Vulkan in Node.js and modern WebGPU browsers.
Every frame of these films is rendered by gitframes from TypeScript in examples/. Click a still to watch it on YouTube.
[!NOTE] Using an AI coding agent? Install the gitframes skills in one line.
Claude Code
/plugin install gitframesAny other agent (Codex, Cursor, Hermes, Gemini CLI, Copilot, and more)
npx skills add gatewai-dev/gitframesSee Agent Skills & Plugins for details.
Modern automated video generation is usually constrained by the architectures of general-purpose web browsers: process overhead, non-deterministic DOM layout reflows, and slow screenshot capture. Gitframes treats video composition as software engineering:
| Pillar | What it means | |
|---|---|---|
| 🚀 | Zero Headless-Browser Overhead | No Puppeteer, no Chromium IPC, no page.screenshot(). Gitframes talks straight to native GPU devices via Dawn/WebGPU and hardware-encodes with @napi-rs/webcodecs. |
| 🎯 | Deterministic Frame-Accurate Clock | Absolute frame clocks, discrete sample points, and frame-accurate audio BeatGrids. No floating timers, no drift, no dropped frames. |
| 🔠 | Analytic, Resolution-Independent Type | The Slug algorithm evaluates glyph contours per-pixel in WGSL — no texture atlases, no scaling artifacts, razor-sharp from 10 px to 10,000 px. |
| 🎨 | Photoshop-Grade Tonal & Spatial VFX | 50+ modular GPU shaders: Curves, Levels, Selective Color, 3D LUTs, Halftone, Film Grain, Unsharp Mask, Mesh Warp, and Screen-Space Relighting. |
| 🧊 | Unified 3D & 2D Depth Compositing | Nest 2D flex/box trees inside 3D homography planes, multiplane rigs, and meshes (OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF), with PBR glass and SSAO. |
| 🔊 | Built-in Procedural Audio DSP | Multi-track soundtracks, deterministic procedural transition SFX (whoosh, impact, riser), and reactive signals that drive visuals from audio. |
| 👁️ | On-Device Neural Vision | Object tracking, instance segmentation, multi-person pose, and person mattes from Apache-2.0 ONNX models — feeding reactive signals without a round trip to disk. |
| ☁️ | Cloud-Native & CI/CD Ready | ~200–400 MB RAM per worker (vs. 2–4 GB for Chromium), ideal for serverless GPU render clusters (AWS G4/G5, Modal, RunPod, Kubernetes). |
Developers generating video programmatically commonly weigh Remotion (React/Chromium) or Hyperframes (Canvas2D/SVG web animation). The matrix below compares the fundamental engineering dimensions.
| Capability / Dimension | Gitframes | Remotion | Hyperframes |
|---|---|---|---|
| Underlying Engine | Native WebGPU (WGSL compute & render pipelines via Dawn / Metal / Vulkan) | Chromium / Puppeteer (React DOM, HTML/CSS layout) | Canvas2D / WebGL / SVG (browser or Node Skia) |
| Rendering Architecture | Direct hardware framebuffer rendering & hardware video encoding (@napi-rs/webcodecs) | Spawns headless Chrome; captures frames via CDP / page.screenshot() | Software or hardware 2D canvas context |
| Throughput | 60–120+ FPS (real-time to faster-than-real-time GPU execution) | 5–20 FPS (DOM reflow, IPC, rasterization) | 20–40 FPS (CPU draw commands / JS) |
| Memory Footprint | ~200–400 MB per render (zero browser) | 1.5–4.0 GB+ per worker (Chromium + V8 DOM heap) | ~500 MB–1 GB (Skia/Canvas bindings) |
| Typography Engine | Slug GPU — analytic Bézier evaluation in WGSL, infinite zoom, After Effects selectors | Browser DOM text (CSS fonts, rasterized, blurry under 3D transforms) | Canvas2D / path text (CPU-rasterized glyphs) |
| 2D VFX & Post-Processing | 50+ WebGPU shaders (Curves, Levels, Selective Color, 3D LUT, Film Grain, Halftone, Liquify, PBR Glass, Relight) | CSS Filters or custom WebGL canvas wrappers | Basic Canvas2D composites and 2D filters |
| 3D Graphics & Depth | Native 3D scene graph — LookAt/Turntable camera, multiplane, skinning (OBJ/FBX/glTF), SSAO, PCSS, DoF | None built-in (embed Three.js/Fiber inside React DOM) | Minimal 2.5D layers; no unified mesh pipeline |
| Motion Blur & Physics | Physical 180° shutter velocity buffers in MRT + closed-form spring kinematics | CSS transitions / JS interpolation; synthetic blur hacks | Frame interpolation or manual multipass |
| Audio Engine & DSP | Native audio DSP & procedural SFX (multi-track mixing, beat grids, reactive signals) | <Audio> playback; basic volume curves | Basic static audio playback |
| Charts & Data Viz | Layer.chart — line, area, bar, scatter, candlestick, pie and donut charts built from native vector nodes, with staggered reveal animations | DOM chart libraries (Recharts, Chart.js) | Custom canvas draw operations |
| AI & Computer Vision | On-device ONNX vision — COCO-80 detection + instance masks (RTMDet-Ins), COCO-17 pose (RTMO), person mattes (Selfie Segmenter); WebGPU tensor conditioning (Canny, depth-to-normals, optical flow, deflicker) | External pre-rendered assets; no native GPU tensor conditioning | External pre-rendered assets |
| Headless Verification | FrameGrid contact sheets, single-frame snapshots, Skia MSE pixel-invariant assertions | Playwright/Puppeteer visual snapshots | Manual frame inspection / canvas diffing |
| Docker / Cloud Portability | Compact (~500 MB slim image with native GPU/Vulkan drivers) | Heavy (~2–3 GB with Chromium, fonts, X11/Mesa) | Moderate container size |
Traditional text relies on CPU rasterization or low-res SDF atlases that soften under 3D camera sweeps. Gitframes integrates the Slug algorithm (SlugPipeline):
square, ramp_up, ramp_down, triangle, smooth), easeHigh/easeLow curves, and seeded PRNG character shuffling (TextAnimator).TypewriterAnimator).evaluateVolumetricFormation).A comprehensive suite of professional image/video shader nodes in nodes/ and packages/webgpu-renderers:
ApplyLUT).PBRGlass).Camera3D) calibrated so z = 0 matches 2D canvas pixel coordinates 1:1.Layer3D.cube, carousel, prism, plane, grid with unified depth-buffer testing.rg16float MRT buffers..audio media nodes with frame-exact lifecycle control.renderSfx, mixSfxInto, softLimit).mixAudioTracks and encodeStereoWav.Signal.builder) or audio analysis.Layer.chart builds line, area, bar (grouped or stacked), scatter, candlestick, pie and donut charts. d3 computes the scales, ticks and geometry; every bar, line, slice and label is an ordinary box, path or text node:
animate: { start, duration, stagger, ease }, or animate: false).Layer.chart(
{
type: "bar",
width: 900,
height: 480,
categories: ["Q1", "Q2", "Q3", "Q4"],
series: [
{ name: "Revenue", data: [12, 19, 24, 31] },
{ name: "Costs", data: [8, 11, 13, 15] },
],
yAxis: { format: "$,.0f" },
animate: { start: 10, duration: 30 },
},
{ position: "absolute", x: 120, y: 200 },
);
@gitframes/vision runs ONNX models via onnxruntime-node (CPU) or onnxruntime-web (WebGPU) and wires every result into the same reactive signal surface the rest of Gitframes consumes.
[!TIP] Lazy by construction.
VisionRunner.create(),comp.withVision(...)andVisionNode.attach(...)perform zero I/O — no downloads, no sessions, no file probes. A model is fetched the first time a task actually runs. To warm up ahead of time, callawait runner.preload(["detect", "pose"])(orawait vision.ready()on an attached node).
Every model is Apache-2.0, pinned to an immutable Hugging Face revision, and verified by SHA-256 after download.
| Task | Option | Model | Output |
|---|---|---|---|
| Detect | enableDetection | RTMDet-Ins t/s/m (OpenMMLab) | COCO-80 boxes + scores, tracked over time |
| Segment | enableSegmentation | RTMDet-Ins (same forward pass as detect) | Soft per-instance masks, frame-aligned |
| Pose | enablePose | RTMO t/s/m (OpenMMLab) | 17 COCO keypoints + visibility per person |
| Matte | enableMatte | MediaPipe Selfie Segmenter (Google) | Fast person-vs-background alpha for portrait / webcam framing |
variant: "t" | "s" | "m" (default "s"; ~24 / 43 / 116 MB for RTMDet-Ins). Tune confidence and a COCO classes filter per composition. On CPU, a 2K frame takes roughly 200–340 ms to detect + segment, ~120 ms for pose and ~20 ms for the matte with "s".matteSource: "instance", the default).mask / matte / crop modes merge every comparably sized instance that overlaps the main subject, so a flowing dress or a held instrument stays attached to the person, while a tunnel or window framing them does not.$GITFRAMES_MODELS_DIR (default ~/.cache/gitframes/models). Point baseUrl or GITFRAMES_MODELS_BASE_URL at your own mirror for air-gapped or CI renders.PoseSkeletonRenderer).SegmentationTexturePool).TemporalObjectTracker) assigns stable trackIds via IoU association, with configurable minHits, positionSmoothing, and velocity-based coasting for up to maxMissedFrames (default 15) so a transient miss holds the track instead of flashing.pose-track-matcher) binds keypoints to the right track by id, then by spatial IoU fallback.comp.analyzeVisionSequence(src, { tasks, categories }) decodes frames through the mediabunny pipeline, tracks them, and returns a zod-serializable report (per-track frame ranges, mean speed, sampled center paths, per-class presence/confidence, mean mask coverage, model download bytes/timing) (analyzeSequence).Every tracked entity is exposed as reactive ProgrammaticSignals that animate layers and shader uniforms:
| Group | Highlights |
|---|---|
objects | get(trackId), byCategory(cat, rank), primary, count, hasCategory, detectedCategories |
objects.*.bounds | x/y/width/height, screenX/screenY/screenWidth/screenHeight, aspectRatio, area |
objects.*.anchors | 9 anchors (corners, edges, center) ready for pinning |
objects.*.kinematics | vx, vy, speed, acceleration, headingRad/Deg |
objects.*.pose | All 17 COCO keypoints, plus hasPose, wristSpeed, handRaised, bodyTiltAngle |
masks | get(trackId), subject, count; per-mask area, coverage, solidity, bboxFill |
segmentation | subject, humanSilhouette, instanceMasks, matte.coverage, GPU stencilTexture |
classes | Per-class count, maxConfidence, present, primary, plus a detection histogram |
| Tensors | poseLandmarksTensor [17,3], objectsTensor [16,8], masksTensor [16,2], histogramTensor [80] |
Project normalized landmarks to screen space with a configurable camera FOV, then bind any node to a track or landmark (SpatialLandmarkTransformer, spatial-pin):
pinToObject(track, { anchor, offsetX/Y/Z, matchWidth, matchHeight, smoothFrames, hideWhenLost })pinToLandmark(coord, { offsetX/Y/Z })comp.addSubjectSandwich({ source, behind, feather, fit }) cuts the foreground subject out and places typography/graphics behind them.comp.addSmartFraming({ source, target, targetAspect, damping, leadHeadroom }) auto-crops 16:9 → 9:16 while tracking target.comp.addSubjectOutline(vision.segmentation.subject, { source, color, width, blur }) strokes the segmented boundary as an audio-reactive contour glow.layer.blurRegion(track, { strength }) blurs faces, plates, or any detected class.passthrough, mask, matte, crop, skeleton, boxes, tracking; pick the cutout alpha with matteSource: "instance" | "selfie", and optionally keyBackground to grow the subject into connected foreground.@gitframes/vision/schemas) so the hot path stays zod-free. Unknown or removed options are rejected, not silently ignored.vision.summary(frame) returns a deterministic, serializable snapshot (objects, classes, masks) safe to call inside a frame hook.@gitframes/vision/web re-exports the engine plus createWebGPUProvider() / hasWebGPU(); onnxruntime-web is an optional lazy peer.skia-canvas to verify shader math, font coverage, and Mean Squared Error (MSE) temporal deltas.comp.renderFrameGrid(...) outputs sequential-frame contact sheets for instant review of easing, kinetic type, and transitions.startPreview({ entry, export }) serves a localhost WebGPU player that loads the composition's own module and renders every frame live in the browser. Nothing is streamed: the server only hands over the bundle, the project's assets, and the soundtrack mixed by the export engine.startPreview serves the page and returns its URL instead of opening a browser, so an agent can show it in its own pane (Claude Code, Codex); pass open: true to open the system browser.import { startPreview } from "gitframes";
const session = await startPreview(
{ entry: new URL("./film.ts", import.meta.url), export: "buildFilm" },
{ title: "gitframes film" },
);
console.log(`Preview at ${session.url}`);
await session.closed; // serves until its tab closes or a newer preview takes over
Managed with pnpm workspaces and turbo:
gitframes/
├── packages/
│ ├── gitframes/ # Unified SDK (Composition, Layer, LayerAnimation, Signal, effects)
│ ├── core/ # Core AST, Effect base class, VirtualMediaData, vision types
│ ├── compositions/ # Layout engine, Flex/Box AST compiler, timeline evaluator
│ ├── webgpu-renderers/ # WGSL shaders, Slug text engine, 3D renderer, camera, lights, materials
│ ├── tensor-webgpu/ # WebGPU compute pipelines (Canny, depth-to-normals, flow, deflicker, landmarks)
│ ├── vision/ # ONNX vision engine: detect, segment, pose, matte, tracking, signals
│ ├── renderer/ # Headless Node.js WebGPU renderer via Dawn, WebCodecs, skia-canvas
│ ├── renderers/ # Higher-level render orchestration
│ ├── media/ # Media decoding / encoding adapters
│ ├── node-sdk/ # Node renderer contracts and result schemas
│ ├── server-utils/ # Server infrastructure, storage, asset caches
│ └── client-utils/ # Shared browser utilities
├── nodes/ # 58+ specialized domain nodes (VFX, audio, layout, node-vision)
├── apps/
│ └── renderer-service/ # Production HTTP / gRPC rendering microservice container
├── examples/ # Reference compositions and films
├── plugins/gitframes/ # Agent plugin: skills only (setup, compose, effects, render)
└── scripts/ # Build, release, and plugin validation tooling
pnpm add gitframes
Requirements: Node.js ≥ 22. Gitframes uses native GPU acceleration via Dawn / WebGPU or Vulkan.
import { Composition, Layer, LayerAnimation } from "gitframes";
// 1. Initialize a 1080p60 composition
const comp = new Composition({
width: 1920,
height: 1080,
fps: 60,
durationFrames: 180, // 3 seconds
backgroundColor: "#090a0f",
fonts: ["assets/fonts/Inter.ttf", "assets/fonts/SpaceGrotesk.ttf"],
});
// 2. Define physical snap-overshoot animations
const cardEntrance = LayerAnimation.create()
.fadeIn(0, 20, "power2.out")
.fromTo("y", 60, 0, { start: 0, end: 35, ease: "back.out(1.5)" })
.fromTo("scale", 0.92, 1.0, { start: 0, end: 35, ease: "back.out(1.2)" });
// 3. Assemble a responsive flex-layout card
const heroCard = Layer.box({
width: 720,
height: 380,
background: "#141721",
borderRadius: 24,
borderColor: "#262b3d",
borderWidth: 1.5,
padding: 32,
children: [
Layer.flex({
dir: "column",
gap: 16,
children: [
Layer.text("GITFRAMES ENGINE", {
fontSize: 16,
fontWeight: 700,
fill: "#6366f1",
letterSpacing: 2.0,
}),
Layer.text("Next-Gen WebGPU Motion", {
fontSize: 48,
fontWeight: 700,
fill: "#f8fafc",
fontFamily: "SpaceGrotesk",
}),
Layer.text("Direct hardware video composition without headless browser overhead.", {
fontSize: 20,
fill: "#94a3b8",
lineHeight: 28,
}),
],
}),
],
}).animate(cardEntrance);
comp.add(heroCard);
import { Composition, Layer, Layer3D, CameraAnimation, Light } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60, durationFrames: 300 });
// 1. LookAt 3D camera with a continuous orbit
const cameraAnim = CameraAnimation.camera().orbit({
azimuth: { from: -30, to: 30 },
elevation: { from: 15, to: 15 },
radius: { to: 1200 },
start: 0,
end: 300,
});
comp.add(
Layer.camera({ x: 960, y: 540, z: -1000, targetX: 960, targetY: 540, targetZ: 0 }).animate(cameraAnim)
);
// 2. Studio lighting
comp.add(Light.ambient("#ffffff", 0.4));
comp.add(Light.directional({ color: "#e0e7ff", intensity: 1.2, x: 500, y: -800, z: -600 }));
// 3. 3D model with skeletal animation
comp.add(
Layer.glb("assets/models/character.glb", {
x: 960,
y: 640,
z: 0,
scale: 2.5,
material: "lit",
loop: true,
})
);
// 4. 3D prism layout carousel
comp.add(
Layer3D.carousel({
radius: 400,
items: [
Layer.box({ width: 280, height: 180, background: "#1e293b", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#334155", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#0f172a", borderRadius: 16 }),
],
})
);
import { Composition, Layer, LayerAnimation, Signal, renderSfx, mixSfxInto, softLimit } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
const totalFrames = 240;
// 1. Soundtrack layer
comp.addAudio(Layer.audio("assets/score.mp3", { volume: 0.9, durationFrames: totalFrames }));
// 2. Frame-accurate procedural SFX on the beat grid
const bed: [Float32Array, Float32Array] = [
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
];
mixSfxInto(bed, [
renderSfx({ type: "whoosh", atBar: 0.79, volume: 0.5 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 1 }),
renderSfx({ type: "impact", atBar: 1.0, volume: 0.8 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 2 }),
]);
softLimit(bed);
// 3. Tempo signal (120 BPM = 2 Hz)
const beatPulse = Signal.builder({ type: "sawtooth", frequency: 2, amplitude: 0.08, offset: 1.0 });
// 4. Bind it to visuals
const reactiveCard = Layer.box({ width: 400, height: 250, background: "#1c202e", borderRadius: 20 })
.animate(
LayerAnimation.create()
.signal("scale", beatPulse, { multiplier: 1.0, offset: 0.0 })
.fromTo("opacity", 0, 1, { start: 0, end: 15, ease: "power2.out" }),
);
comp.add(reactiveCard);
import { Composition, FilmGrain, Vignette, ColorBalance } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
// Whole-composition cinematic grade + film emulsion
comp.apply(new Vignette({ strength: 0.28, radius: 0.85 }));
comp.apply(new FilmGrain({ strength: 0.06, size: 1.5, animated: true }));
comp.apply(
new ColorBalance({
shadows: { cyanRed: 0, magentaGreen: 2, yellowBlue: 6 },
highlights: { cyanRed: 4, magentaGreen: 1, yellowBlue: -2 },
}),
);
import { Composition, Layer, Vignette } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 30 });
// Run vision on the whole composition. Models download lazily on first use.
const vision = comp.withVision({
enableDetection: true,
enableSegmentation: true,
enablePose: true,
variant: "s",
confidence: 0.35,
});
// Pin a caption to the primary tracked subject (smoothing + auto-hide when lost)
comp.add(
Layer.text("SUBJECT 01", { fontSize: 40, fill: "#f8fafc" }).pinToObject(
vision.objects.primary,
{ anchor: "topCenter", offsetY: -48, smoothFrames: 5, hideWhenLost: true },
),
);
// Drive a shader uniform from a reactive signal — here, subject mask coverage
comp.add(
Layer.box({ width: 1920, height: 1080, background: "#000000" }).withEffect(
new Vignette({ strength: vision.segmentation.subject.coverage, radius: 0.9 }),
),
);
// Or use the one-liners for the common editorial moves:
// comp.addSubjectSandwich({ source: "assets/dancer.mp4", behind: [headline], feather: 4 });
// comp.addSmartFraming({ source: "assets/action.mp4", target: vision.objects.primary, targetAspect: 9 / 16 });
// comp.addSubjectOutline(vision.segmentation.subject, { source: "assets/character.mp4", color: "#FF5A1F", width: 6 });
// Inspect a source before authoring: one-shot, ffmpeg-free, zod-serializable report
const report = await comp.analyzeVisionSequence("assets/street.mp4", {
tasks: ["detect", "pose"],
categories: ["person"],
});
console.log(report.tracks.map((t) => `${t.category}#${t.trackId} ${t.frames.join("–")}`));
Standalone runner (no composition):
import { VisionRunner } from "@gitframes/vision";
const runner = VisionRunner.create({ variant: "s", confidence: 0.3 }); // zero I/O
const frame = { data: rgba, width: 1920, height: 1080 };
const boxes = await runner.detect(frame); // downloads RTMDet-Ins on first call
const { masks } = await runner.segment(frame); // same forward pass, no second inference
const { people } = await runner.pose(frame); // RTMO, COCO-17 keypoints
runner.close();
In the browser (WebGPU EP):
import { VisionRunner, createWebGPUProvider, hasWebGPU } from "@gitframes/vision/web";
if (hasWebGPU()) {
const runner = VisionRunner.create({ provider: createWebGPUProvider() });
}
import { buildMyComposition } from "./my-composition.js";
const comp = await buildMyComposition();
// 1. Single frame to a PNG buffer for visual inspection
const frameBuffer = await comp.renderFrame({ frame: 45 });
// 2. Contact-sheet grid of 12 sequential frames
const gridBuffer = await comp.renderFrameGrid({
startFrame: 0,
endFrame: 120,
stepFrames: 10,
cellWidth: 320,
showLabels: true,
});
// 3. Final hardware-encoded MP4 with mixed audio
const { filePath } = await comp.renderVideo({
outputPath: "output/final-product-film.mp4",
quality: "high",
concurrency: 4,
});
console.log(`Video rendered successfully to: ${filePath}`);
THEME for colors, type, radii, and spacing. Never hardcode magic hex values or ad-hoc margins.color * opacity * alpha) must use srcFactor: "one" in their blend state ({ srcFactor: "one", dstFactor: "one-minus-src-alpha", operation: "add" }). Never use srcFactor: "src-alpha" for premultiplied output — squaring alpha darkens fades into murky gray.back.out(1.4–1.7) for snap-overshoot entrances, spring / expo.out for decelerating motion, power2.in for exits. Reserve linear for infinite spinners and time counters.skia-canvas pixel sampling in Vitest before shipping.Gitframes ships agent skills that teach Claude, Codex, and other coding agents how to write, render, and check compositions. The plugin (gitframes) is listed in Anthropic's official plugin directory and contains only skills — no MCP servers, hooks, or commands. Every other agent gets the same skills through the skills CLI.
| Skill | Use it for |
|---|---|
gitframes | Starting a project: install from npm, scaffold a composition and render script, first verified render |
gitframes-compose | Compositions, layer trees, layout, animation and easing, beat grids, film structure |
gitframes-effects | Effect classes, the unified section architecture, premultiplied-alpha invariants, vision conditioning |
gitframes-render | Headless rendering, FrameGrid inspection, pixel probes, MP4 delivery checks |
Once installed, skills load automatically when a task matches (e.g. "add a film-grain pass to this scene" or "render a frame grid of intro.ts").
The plugin is instructions only. It bundles no executables, MCP servers, hooks, or package launchers, and it sends no data anywhere. The skills tell your agent to add the gitframes npm package to your project and how to use it. When that code uses on-device vision, the SDK downloads the pinned model weights from Hugging Face on first use (see On-Device Vision). Nothing else leaves your machine.
/plugin install gitframes
Or from your shell:
claude plugin install gitframes@claude-plugins-official
It installs from Anthropic's official marketplace, which Claude Code adds for you, so there is no marketplace step, and plugins from it update automatically. Afterwards, restart Claude Code or run /reload-plugins. /plugin commands need an interactive claude terminal; in the desktop app's Code tab, use the shell form or + > Plugins > Add plugin and pick Gitframes.
Add --scope project to the shell form to record the plugin in .claude/settings.json for the whole team.
Enable it for everyone in your repo. Commit this to .claude/settings.json; Claude Code prompts teammates to install it when they trust the folder:
{
"enabledPlugins": {
"gitframes@claude-plugins-official": true
}
}
Straight from this repository (tracks main instead of the directory release):
/plugin marketplace add gatewai-dev/gitframes
/plugin install gitframes@gitframes-plugins
The skills CLI installs the skills into any of 70+ agents, including Codex, Cursor, Hermes, Gemini CLI, GitHub Copilot, Windsurf, OpenCode, and Goose:
npx skills add gatewai-dev/gitframes
It detects the agents on your machine and asks where to install. To choose them yourself, pass -a once per agent, add -g to install for your user instead of this project, and -y to skip the prompts:
npx skills add gatewai-dev/gitframes -a codex -a cursor -a hermes-agent -g -y
Keep them current with npx skills update, and remove them with npx skills remove.
Or copy the folders by hand: put plugins/gitframes/skills/<name>/ into .claude/skills/, .agents/skills/, or ~/.agents/skills/. VS Code / Copilot / Cursor / Kiro can load the portable plugin.json through their plugin UI.
The plugin lives in plugins/gitframes/ so installs carry only the skills; users get the engine from npm. Two manifests there describe it: plugin.json (portable Agent Plugins 1.0, which also carries the OpenAI listing metadata) and .claude-plugin/plugin.json. The marketplace catalog is .claude-plugin/marketplace.json. The portable field set is closed — client-specific fields go in that client's manifest, not in plugin.json. The version in both follows the gitframes package: pnpm run version:packages syncs it after changeset version (or run pnpm run sync:plugin-version on its own), since clients use it to decide when to update.
Inside this repository, Claude Code and other agents pick up skills through the symlinks in .agents/skills/ and .claude/skills/. Skills live only under plugins/gitframes/skills/; never copy them elsewhere. pnpm run check:plugins validates manifests, skill frontmatter, marketplace catalogs, symlinks, and the generated effects catalog. pnpm run sync:effects-catalog regenerates the gitframes-effects catalog after any Effect class change.
The examples/ directory holds production-grade reference compositions:
| Example | What it demonstrates |
|---|---|
19_gitframes_film | The 30-second master brand film — full pipeline, audio, VFX, 3D |
21_full_circle | Multi-scene narrative composition |
22_gitframes_launch | Launch/product-motion composition |
Gitframes uses pnpm (10+) and turbo for orchestration.
# Install
pnpm install
# Build all packages
pnpm build
# Run conformance tests
pnpm test
# Check the vision models end to end (downloads ~380 MB of weights once)
pnpm --filter @gitframes/vision test:models
# Render a specific showcase example
pnpm --filter @gitframes/example-21-full-circle render
# Render the master brand film
cd examples/19_gitframes_film && pnpm render
An optimized Dockerfile.renderer deploys the renderer service into cloud GPU clusters:
docker build -t gitframes-renderer -f Dockerfile.renderer .
Gitframes is open-source software licensed under Apache-2.0. The vision models it downloads on demand — RTMDet-Ins and RTMO (OpenMMLab) and the Selfie Segmenter (Google) — are also Apache-2.0; see registry.ts for exact sources and checksums.