ai-video-editor is a free, open source photo & video editors project written in JavaScript and released under MIT. It has 859 GitHub stars, 108 forks and 7 open issues, and was last pushed 3 days ago. On this registry it ranks #39 of 41 tracked projects in Photo & Video Editors, with 5 head-to-head comparisons available.

What is ai-video-editor?

Timeline Studio, published on GitHub as ai-video-editor, is an MIT-licensed, local-first video editor that runs in the browser, built for creators and AI agent developers who edit the same real multi-track timeline.

What it is

Timeline Studio is a JavaScript browser application — the topics list React, PWA, browser-ai and ONNX — that pairs a CapCut-style multi-track timeline with AI features executed locally: WebGPU-driven AI music and AI repair, multilingual voiceovers, automatic captions, talking-avatar generation, and deterministic offline export. The interface and its messages are maintained in 13 languages, including Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Thai, Vietnamese and Russian. The repository carries 859 stars, 108 forks and 7 open issues, sits in the Miscellaneous / Photo & Video Editors category, and shows a dense development log through September 2026.

The concrete problem it solves is the split workflow that AI-assisted editing usually forces. A creator who wants generated music, timed watermark removal, voiceover, captions or a talking avatar has generally had to move footage out of the editor and into separate hosted services, then bring the result back. Timeline Studio puts those capabilities behind the same timeline used for manual work such as trimming, overlay transforms and caption styling, and it runs local-first so the media stays in the browser. It also exposes that timeline to software agents: through WebMCP, 21 browser tools cover 26 reviewed edits, allowing an agent to move timed clips, trim sources, change overlay transforms, and request rendered frame or audio samples before anything is inserted. The named model stack includes JoyVASA, LivePortrait, ModelScope and ONNX, and the project enforces a responsible-use policy covering deep-synthesis output.

Key capabilities

  • A CapCut-style multi-track timeline whose white playhead previews the picture live while dragged, snaps to markers and range edges with a shared alignment guide and time readout, and scrubs freely while Alt is held.
  • WebGPU AI music and AI repair, including timed watermark removal presented with a before-and-after review step.
  • Multilingual voiceover and automatic captions generated in the browser, with progress reporting and cancellation.
  • Talking-avatar generation, with JoyVASA, LivePortrait, ModelScope and ONNX named in the topic list as the model stack.
  • WebMCP agent tooling spanning 21 browser tools and 26 reviewed edits, covering timed-clip movement, source trimming, overlay transforms, caption style and position, and project framing.
  • Audio export and pause removal: trimmed clips and the complete timeline mix export as WAV or MP3 with playback speed, volume, fades and space effects applied, while Smart uses Silero VAD to detect long speech pauses for review, ripple editing and undo.
  • Deterministic offline export, where failed exports keep the error visible and can be retried with the same settings.

Who uses it and how

  • Individual creators who want a CapCut-style editing surface without installing a desktop application, using either the hosted editor or the Hugging Face Space at haixin/timeline-studio.
  • Agent developers building editing skills: the repository carries the agent-skills topic, a skills.sh listing, and the WebMCP tool surface, while the companion AI Video Editing Skills Handbook publishes reproducible before-and-after recipes.
  • Multilingual editing and review workflows, where UI copy, controls and messages are maintained across 13 interface languages.
  • Repair and deep-synthesis work such as timed watermark removal with before-and-after review, voiceover, and Smart pause removal with adjustable thresholds, retained gaps and cancellation.
  • Contributors and testers: pull requests are welcomed through CONTRIBUTING.md, with a public Roadmap file, a Releases page and an Issues tracker used for focused tasks and bugs.

Getting started

The README points to a hosted build rather than an installable package: the live editor at video-editor.ai-creator.top, with a Hugging Face Space offered as an alternative. No npm package, Docker image or compose file is named in the README; the project runs as a browser-based, local-first application.

How it compares

The only comparable product named in the project's own materials is CapCut, whose multi-track timeline the editor deliberately mirrors in behaviour and layout. The differences the facts support are licence and locality: Timeline Studio is MIT-licensed, runs inside the browser, and exposes that same timeline to AI agents through 21 WebMCP browser tools rather than treating agent edits as a separate layer.

When to use it — and when not to

Because it is local-first and browser-based, the README names no server-side database, storage tier or SMTP requirement, so the practical prerequisite is a browser capable of running WebGPU inference. Anyone who needs the deep-synthesis features should read the responsible-use section first: the tool is stated as intended solely for technical research and learning, requires lawful authorization for any facial images or videos used, and places all resulting liability on the user. The documentation is also uneven — the README reads largely as an engineering changelog and defers the walkthrough to an external Skills Handbook and Roadmap, so users wanting a single complete manual may find it thin.

project readme (upstream, from github) — read inline

Timeline Studio — Browser AI Video Editor

English | 中文 | 日本語 | 한국어 | Español | Français | Deutsch | Português | ไทย | Tiếng Việt | Русский

Live Demo MIT License skills.sh PRs Welcome LINUX DO Timeline Studio - Listed on Tool Index

Responsible use of deep synthesis

This tool uses deep-synthesis technology and is intended solely for technical research and learning.

Users must ensure that they:

  • use only facial images or videos of themselves or people who have provided lawful authorization;
  • do not create or distribute any illegal, infringing, false, or misleading content;
  • do not present generated content as authentic footage or impersonate another person without their consent.

Users are solely responsible for any legal liability arising from violations of these requirements.

Project updates

  • 2026-09-21 — Remove pauses: Smart now detects long speech pauses locally with Silero VAD, lets you review and select cuts, and trims the selected main-track video with its source audio. Adjustable pause thresholds, retained gaps, ripple editing, cancellation and undo are included, with direct UI copy in all 13 languages.
  • 2026-09-20 — Playhead marker snapping: Dragging the white playhead previews the picture live and snaps to markers and range edges, with a shared alignment guide and time readout. Hold Alt to scrub freely. Marker details show Done until edited, then Apply changes.
  • 2026-09-17 — WebMCP finishing and review: 21 browser tools now cover 26 reviewed edits, including timed-clip movement and source trimming, overlay transforms, caption style/position and project framing. Agents can request rendered frame/audio samples, run browser-local voiceover or transcription with progress and cancellation, and review results before timeline insertion. New controls and messages support all 13 interface languages.
  • 2026-09-15 — Edited audio export: exporting an audio clip now renders its trimmed range with playback speed, volume, fades and space effects applied. Clips and the complete timeline mix can be exported as WAV or MP3; audio-only export is available alongside video export.
  • 2026-09-14 — Export reliability: clips with solid-color backgrounds and opacity keyframes now export correctly without a person mask. Failed exports keep the error visible and can be retried with the same settings.

See the public Roadmap for planned work, Releases for shipped changes, and Issues for focused tasks and bugs.

What can it produce?

Explore reproducible before/after examples and editing recipes:

AI Video Editing Skills Handbook

MartinDelophy%2Fai-video-editor | Trendshift MartinDelophy%2Fai-video-editor | Trendshift

Timeline Studio is a local-first AI video editor that runs in the browser. It combines a CapCut-style multi-track timeline with WebGPU AI music and repair, multilingual voiceovers, automatic captions, talking-avatar generation, and deterministic offline export.

Open the editor · Watch on YouTube · Hugging Face Space

Video demo

Auto Edit

Visual analysis, keyframes, captions, and export.

https://github.com/user-attachments/assets/e8327caa-429e-40ff-a7fe-a59e6cf7a464

AI Repair

Timed watermark removal with before/after review.

https://github.com/user-attachments/assets/aea9f5b4-c720-4b0c-9067-5ec124eef982

AI Voiceover

https://github.com/user-attachments/assets/304a744e-d620-4380-9c17-19af3726f5a4

Timeline Studio editor

AI capabilities

  • Multilingual voiceover: Chinese and mixed Chinese/English Hojo TTS Light 80M FP16 reference voices (晴岚 and 若溪), with autoregressive generation on WebGPU and stable waveform decoding on WASM; English Kokoro 82M; and browser Piper voices for German, Spanish, French, Italian, and Brazilian Portuguese.
  • Local AI music: Stable Audio 3 Small Q4 ONNX runs through WebGPU with translated free-form prompts, 30/60/90/120-second choices, waveform-aware long-track looping, persistent model caching, and automatic insertion into My Assets.
  • Automatic captions: Whisper small q8 ONNX with waveform-aware timing and conservative Chinese recognition cleanup.
  • Smart framing: YOLOS tiny subject detection and MODNet portrait matting for smart crop, caption avoidance, and background removal across images and complete videos.
  • AI Repair: browser-local MI-GAN watermark/object removal with multiple timed repair regions, plus NanoVSR 644K WebGPU 4× restoration for images and videos with synchronized before/after comparison.
  • AI vocal separation: isolate vocals and place the instrumental stem on the music track without leaving the browser workflow.
  • Digital human: JoyVASA audio-to-motion and LivePortrait neural rendering with WebGPU, 256px preview and 512px quality paths.
  • Local-first inference: large models are lazy-loaded, revision-pinned, and cached by the service worker; supported workflows run without uploading project media to an editing backend.

Resilient model delivery

Browser AI models are mirrored on both Hugging Face and ModelScope. Timeline Studio uses ModelScope first for Chinese interfaces and domestic sessions, uses Hugging Face first elsewhere, remembers the first working source for the current runtime, and automatically falls back to the other source when needed. Cache identities are shared across both providers, so switching routes does not download the same revision twice.

Editing and export

  • Contiguous main Visuals track plus timed picture-in-picture overlays.
  • Direct canvas selection, movement, proportional resize, rotation, masks, filters, effects, animation, speed, and explicit keyframes, plus desktop four-way color grading with fully keyframeable temperature, tint, saturation, hue, wheel saturation, and luminance controls.
  • Captions, stickers, voiceover, separated source audio, and music on independent timed tracks.
  • CapCut-style snapping, alignment guides, clip menus, split/duplicate/delete, timeline zoom, undo/redo, and portable .timeline projects.
  • Native media playback for a responsive preview; export uses a separate deterministic offline rendering path.
  • WebCodecs MP4/WebM composition with shared preview/export geometry, audio mixing, captions, overlays, effects, and MediaRecorder fallback.
  • Installable PWA with a cached app shell and multilingual UI.

Agent Skill

The repository includes the AI Video Editing Skill for Codex, Claude Code, Copilot and Gemini CLI, backed by edit-timeline-studio for planning, executing, and verifying editable video timelines.

It helps an agent:

  • inspect media and preserve the user's editing brief;
  • describe reversible edits with stable clip IDs and explicit timestamps;
  • operate the hosted or local editor through the browser compatibility path;
  • validate declarative edit plans with skills/edit-timeline-studio/scripts/validate_edit_plan.mjs;
  • verify track placement, transitions, captions, overlays, audible audio, and final export artifacts;
  • keep the editable .timeline project as the source of truth instead of returning only an opaque render.

The versioned headless command runner loads and inspects portable projects, validates revisioned JSON plans, applies supported operations transactionally, supports dry runs and idempotent operation IDs, and writes a new .timeline archive without rewriting its media files. It also renders the documented portable Visuals + Voiceover + Music subset to a verified H.264/AAC MP4. Browser control remains the compatibility path for operations that are not in the command registry yet. The browser WebMCP integration provides 21 tools for the open project: structured inspection, 26 reviewed visual/caption/audio/marker/framing operations, existing-asset and picture-in-picture insertion, rendered frame/audio sampling, local voiceover/transcription jobs, guarded undo, project saving, and real video export with prepared settings, progress, output receipts and cancellation.

npm run agent -- project.inspect /absolute/path/project.timeline
npm run agent -- track.inspect /absolute/path/project.timeline visuals
npm run agent -- clip.inspect /absolute/path/project.timeline visual-123
npm run agent -- transcript.inspect /absolute/path/project.timeline voice-123
npm run agent -- project.diff /absolute/path/edit-plan.json
npm run agent -- project.run /absolute/path/edit-plan.json
npm run agent -- project.render /absolute/path/render-request.json

The legacy inspect and run aliases remain available. The write registry supports ffprobe-backed, hashed visual/audio import to Visuals, Music, or the portable Voiceover slot, plus transactional timed edits, captions, source-accurate Visuals/Overlays, transitions, validated properties, track state, and ratio changes. Read commands return project, track, clip, transcript, media-inventory, and field-level predicted diffs. See the command contract for the plan envelope.

Install through the public skills.sh directory (the current CLI requires Node.js 22.20.0 or later):

npx skills add MartinDelophy/ai-video-editor --skill edit-timeline-studio

Or install it with GitHub CLI 2.90.0 or later:

# Claude Code
gh skill install MartinDelophy/ai-video-editor edit-timeline-studio --agent claude-code --scope user

# Codex
gh skill install MartinDelophy/ai-video-editor edit-timeline-studio --agent codex --scope user

To install the tested release instead of following the latest release, add --pin v1.0.8. Preview the Skill before installing with:

gh skill preview MartinDelophy/ai-video-editor edit-timeline-studio

Roadmap

  • Now: expand the versioned command registry, harden deterministic offline export, and improve timeline editing reliability.
  • Next: expand headless render parity and the reviewed browser WebMCP command subset.
  • Later: add collaborative review workflows, a plugin extension surface, and more locally verified AI models.

Roadmap priorities are shaped in GitHub Discussions. Feature requests and real-world workflow feedback are welcome.

Help wanted

Timeline Studio is looking for contributors interested in browser media, WebCodecs, WebGPU/ONNX, timeline UX, localization, testing, and documentation.

  • Try the live editor and report reproducible bugs in Issues.
  • Join Discussions to propose features, share projects, or help prioritize the roadmap.
  • Read the contribution guide for setup, validation, and first-contribution guidance.
  • Contributions of focused fixes, tests, translations, documentation, and example projects are especially useful.

Quick start

Requirements: Node.js 20+ and a modern Chromium browser. WebGPU is recommended for the heaviest AI workflows.

git clone https://github.com/MartinDelophy/ai-video-editor.git
cd ai-video-editor
npm install
npm run dev

Open the local URL printed by Vite. The first AI run may download model files; later runs reuse the browser cache.

Validate and build

npm run build
npm run preview

Run the complete repository check with:

npm run check

Deploy

The included netlify.toml builds with npm run build, publishes dist, enables the cross-origin isolation headers required by browser AI/media workers, and provides the SPA fallback.

npx netlify-cli deploy --prod --dir=dist

Support and feedback

If this project helps you, please consider giving it a ⭐ Star. If you encounter a problem, please open an Issue.

Join our Discord community to ask questions, share feedback, and connect with other users and contributors.

License

Timeline Studio's original source code is licensed under the MIT License.

The MIT License does not automatically apply to third-party models, model weights, datasets, fonts, stock media, or other bundled or remotely downloaded assets. Those materials remain subject to their respective upstream licenses and terms, regardless of where they are hosted or how they are downloaded or integrated. Review MODEL_LICENSES.md before redistribution or commercial use.

Frequently asked questions

Is ai-video-editor free to use?

ai-video-editor is open source under the MIT licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does ai-video-editor do?

Open-source, local-first video editor where creators and AI agents edit the same real timeline.

What is ai-video-editor written in?

ai-video-editor is primarily written in JavaScript. Its source is publicly available at https://github.com/MartinDelophy/ai-video-editor, and it has 859 GitHub stars.