arvideo
Sender Director Factory Spotter Download
FACTORY — SELF-HOSTED EDIT LINE

Factory edits the pile
into finished videos.

Factory reads every clip on your drives into a database, then cuts event footage into finished videos with music — on your hardware, on your schedule. Batch mode works through the whole archive on its own, overnight if that's when your machine is free.

Coming soon SELF-HOSTED · $249
HOW A RUN GOES:CUT auto · ai-edit · review MODE standard · guide RENDER normal · overlay ± boxes · star · + diagnostic hud SEE IT RUN →
WHAT COMES OUT

A finished video, per event.

ONE EVENT IN, ONE FINISHED MP4 OUT — MUSIC, TITLE CARD AND ALL.
WHAT IT READS FIRST

Know your footage.

Before Factory edits anything, it reads every clip and writes down what it finds: quality scores, duration, resolution, which camera shot it, and GPS and flight telemetry when the file — or the sidecar next to it — carries it: GoPro's GPMF lives in the file, DJI's SRT beside it. Claude writes a short description of each scene, and Whisper transcribes the audio locally with timestamps. Everything goes into one database, and every edit afterward is built from those records.

SCOREDsharpness · brightness · motion · audio
LOGGEDcamera · duration · resolution · GPMF + SRT telemetry · GPS
DESCRIBEDscene notes by Claude, per clip
TRANSCRIBEDWhisper, timestamped, fully local
DETECTEDoptional — object detection, landmark detection, star detection
PROVENANCEprovenance and rights — every source segment and a generation count, timestamped and written into the finished video's metadata in a tag chosen to survive trims and re-encodes; the most restrictive source wins
CLIP_0834.MP4 GoPro Hero 11 Black · 1080p · 00:02:41
IN THE FILE
camera body duration · resolution timestamps GPS position GPMF + SRT telemetry
EXTRACTED BY AI
scene notes — Claude transcript — Whisper quality scores detections — optional
ONE DATABASE every edit is built from these records
ONE CLIP, READ INTO THE DATABASE.
HOW IT CUTS

Two ways to cut.

AUTO RULES-BASED

Deterministic edits from the index: best takes by quality score, trimmed and sequenced by rule. AI picks the song to fit the event; when the track has a clear structure, cuts snap to the beat.

BEAT SNAPPING AI SONG SELECTION QUALITY FILTERING
AI-EDIT CONTEXT-DRIVEN

AI reads everything the index knows — scene notes, transcripts, scores, telemetry — and writes the cut plan itself: what to keep, in what order, and why. It works from a written editing style guide, not a pile of if-statements, and dissolves are budgeted: at most one per 90 seconds, only across a real time jump.

FULL-CONTEXT CUT PLAN STORY-AWARE ORDERING
CUT
0834
0512
0221
0227
0771
0790
SONG
INTRO
VERSE
CHORUS
VERSE
OUTRO
CUT ON PHRASE BOUNDARY CROSSFADE CLIP EDGES SNAPPED TO THE SONG'S PHRASES
THE MUSICIT CHOOSES, NEVER GENERATES

Claude picks from your own music library, using the event's location, activity, subjects and dialogue — and writes down one sentence of reasoning per pick. Songs are assigned across the whole batch in one pass so an overnight run doesn't repeat itself, and recently used songs sit out a cooldown. Every song is beat-analyzed locally: cuts land on the beat and never straddle a musical phrase boundary, and the music ducks under actual speech — driven by the transcript's word timings, not by which clip is playing.

THE PROMPT — ASK THE WHOLE ARCHIVE Prompt mode queries the whole indexed archive across years: "Austin skateboarding 2020–2022" builds a custom montage, chaining songs when the ask outruns any single track.
THE RECEIPTS — BRIEF + FCPXML Every event ships with its edit receipts: a written brief, the clip table with rejection reasons, and a DaVinci Resolve timeline (FCPXML) — so a human can take over the cut. The machine does the assembly; you keep the final word.
THE SLATE — JOBS, LOGS, AND A BIG RED STOP Every run is a job on the slate: you can watch the log live, stop it cold, or requeue it.
OPT-IN, PER EVENT

The optional passes.

Each of these is a separate pass you turn on when an event deserves it. What they find goes into the index and onto the overlay renders.

DETECTION01

Local models find and follow every subject on your own GPU — only cropped snips of the sharpest frames ever leave the machine, and the subject is identified from those snips. Then three AIs research the subject independently, and Claude checks their answers and keeps only the most true and most interesting ones.

TRK 12 · 0.88
NAF N3N-3 “YELLOW PERIL”997 BUILT
Naval Aircraft Factory · 1935–1942

Float-equipped N3Ns were still flying midshipmen at the Naval Academy in 1961 — the last biplanes in United States military service, a decade into the jet age.

STARS02

A night sky is not a matter of opinion, so star frames are plate-solved against the catalogues — Hipparcos, IAU, Skyfield — instead of asked to any model. What was actually overhead that night comes back labelled.

GROUND03

Flight footage gets its ground truth from the map instead of a model: the telemetry track is queried against OpenStreetMap. The ridge, the reservoir and the road you followed come back named.

DIAGNOSTIC HUD04

The diagnostic render lays the index back over the picture — scores, boxes, track ids, the fact card, and the transcript tracking the timecode word by word. It's the render you watch to check the work, not the one you send out.

RUNS ON YOUR IRON

Requirements.

SERVERa Linux machine on your network hosts it, pointed at your drives
CLIENTthe GUI is a web app in any browser — phone-friendly, so you can run it remotely (Adam drives his over Tailscale)
GPUan NVIDIA GPU is required — the renderer is NVENC-only with no CPU fallback, and transcription runs on CUDA
LOCAL MODELSWhisper transcription · local detection models (segmentation + tracking), on your own GPU
API KEYSbring your own — an Anthropic API key is required to run at all; the detection and fun-fact passes call other models through OpenRouter, and place names come from OpenStreetMap
YOUR DATAyour footage never leaves the machine — only small still frames and transcript text go to the models, and the finished video is never uploaded anywhere. No accounts, no telemetry, and every API call's cost is logged locally, with a per-run cost report
OUTPUT1920×1080, 30 fps, H.264 + AAC MP4, with a five-second title card carrying the event name and date; vertical phone clips get a blurred-background fill instead of black bars
360 FOOTAGErecognized and deliberately excluded from the edit pool — spheres go through Director (or a reframe step) first
FACTORY
Automated batch video editing, object detection and multiple render options.
Coming Soon Want it early? Write to info@arvideo.io
SELF-HOSTED · YOUR DRIVES, YOUR KEYS · NO SUBSCRIPTION
arvideo
SENDER · DIRECTOR · SPOTTER
INFO@ARVIDEO.IO