Vibecode diclip
track this build5 steps, step by step0%The core loop is genuinely one-shottable: transcribe, rank the strongest moments, cut vertical clips with burned captions. The gaps are execution and infra. Face-aware reframing with seat tracking, a full in-browser timeline editor, and a container-scale render pipeline are not a one-session build. A competent agent can produce a usable personal clipper quickly, but matching diclip's reframe quality and editing depth is a multi-day project.
You are building a lean indie version of diclip. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # diclip indie build ## Goal Build the smallest trustworthy replacement for the core diclip workflow for one developer or a tiny team. ## Scope Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - face-aware vertical reframing with seat tracking - full in-browser timeline editor with scenes, tracks, and effects - container-scale render pipeline and yt-dlp import - credit-based capacity model and priority processing - ready-to-post hooks and confidence scoring polished for clips If those capabilities are essential, use diclip instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a personal substitute for diclip, not a platform clone. - Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript frontend, faster-whisper for word-level transcription, and FFmpeg for rendering. - I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the video with speaker timing. - A ranking step selects the strongest 30-90 second moments: call an LLM API key from .env with the transcript, or fall back to a local heuristic (keyword density, question marks, pauses) when no key is set. - Show each candidate clip with its hook line, a confidence score, and a one sentence reason, marked ready or pending. - Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1, or 16:9 output. - Render clips with FFmpeg using center-crop or blur-pad, burning captions styled by speaker. Never overwrite the source. Show progress and a useful failure message. - Store projects as JSON under ~/diclipDIY/projects and renders under ~/diclipDIY/exports, with a recent-projects page and a delete action. - Bind to localhost only. No accounts, telemetry, or network calls after the model and transcript are local, except the optional LLM ranking call. - Deliberately exclude face-aware reframing with seat tracking, the full in-browser timeline editor, cloud container rendering, yt-dlp downloads, and credit billing. - Add unit tests for the ranking fallback and caption grouping, plus one smoke test that imports a short fixture and produces a playable MP4. - Include a README with setup, data paths, and an honest note that CPU transcription and rendering are slow on long videos. ## Required capabilities - FFmpeg - faster-whisper - LLM API key for moment ranking (optional, falls back to heuristics) - desktop with enough storage for source media and renders ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of diclip. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # diclip indie build ## Goal Build the smallest trustworthy replacement for the core diclip workflow for one developer or a tiny team. ## Scope Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - face-aware vertical reframing with seat tracking - full in-browser timeline editor with scenes, tracks, and effects - container-scale render pipeline and yt-dlp import - credit-based capacity model and priority processing - ready-to-post hooks and confidence scoring polished for clips If those capabilities are essential, use diclip instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a personal substitute for diclip, not a platform clone. - Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript frontend, faster-whisper for word-level transcription, and FFmpeg for rendering. - I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the video with speaker timing. - A ranking step selects the strongest 30-90 second moments: call an LLM API key from .env with the transcript, or fall back to a local heuristic (keyword density, question marks, pauses) when no key is set. - Show each candidate clip with its hook line, a confidence score, and a one sentence reason, marked ready or pending. - Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1, or 16:9 output. - Render clips with FFmpeg using center-crop or blur-pad, burning captions styled by speaker. Never overwrite the source. Show progress and a useful failure message. - Store projects as JSON under ~/diclipDIY/projects and renders under ~/diclipDIY/exports, with a recent-projects page and a delete action. - Bind to localhost only. No accounts, telemetry, or network calls after the model and transcript are local, except the optional LLM ranking call. - Deliberately exclude face-aware reframing with seat tracking, the full in-browser timeline editor, cloud container rendering, yt-dlp downloads, and credit billing. - Add unit tests for the ranking fallback and caption grouping, plus one smoke test that imports a short fixture and produces a playable MP4. - Include a README with setup, data paths, and an honest note that CPU transcription and rendering are slow on long videos. ## Required capabilities - FFmpeg - faster-whisper - LLM API key for moment ranking (optional, falls back to heuristics) - desktop with enough storage for source media and renders ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of diclip. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # diclip product brief ## Problem The core loop is genuinely one-shottable: transcribe, rank the strongest moments, cut vertical clips with burned captions. The gaps are execution and infra. Face-aware reframing with seat tracking, a full in-browser timeline editor, and a container-scale render pipeline are not a one-session build. A competent agent can produce a usable personal clipper quickly, but matching diclip's reframe quality and editing depth is a multi-day project. ## Product outcome Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - FFmpeg - faster-whisper - LLM API key for moment ranking (optional, falls back to heuristics) - desktop with enough storage for source media and renders ## Explicit non-goals for v1 - face-aware vertical reframing with seat tracking - full in-browser timeline editor with scenes, tracks, and effects - container-scale render pipeline and yt-dlp import - credit-based capacity model and priority processing - ready-to-post hooks and confidence scoring polished for clips ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build me a personal substitute for diclip, not a platform clone. - Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript frontend, faster-whisper for word-level transcription, and FFmpeg for rendering. - I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the video with speaker timing. - A ranking step selects the strongest 30-90 second moments: call an LLM API key from .env with the transcript, or fall back to a local heuristic (keyword density, question marks, pauses) when no key is set. - Show each candidate clip with its hook line, a confidence score, and a one sentence reason, marked ready or pending. - Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1, or 16:9 output. - Render clips with FFmpeg using center-crop or blur-pad, burning captions styled by speaker. Never overwrite the source. Show progress and a useful failure message. - Store projects as JSON under ~/diclipDIY/projects and renders under ~/diclipDIY/exports, with a recent-projects page and a delete action. - Bind to localhost only. No accounts, telemetry, or network calls after the model and transcript are local, except the optional LLM ranking call. - Deliberately exclude face-aware reframing with seat tracking, the full in-browser timeline editor, cloud container rendering, yt-dlp downloads, and credit billing. - Add unit tests for the ranking fallback and caption grouping, plus one smoke test that imports a short fixture and produces a playable MP4. - Include a README with setup, data paths, and an honest note that CPU transcription and rendering are slow on long videos. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted diclip capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# diclip indie build ## Goal Build the smallest trustworthy replacement for the core diclip workflow for one developer or a tiny team. ## Scope Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - face-aware vertical reframing with seat tracking - full in-browser timeline editor with scenes, tracks, and effects - container-scale render pipeline and yt-dlp import - credit-based capacity model and priority processing - ready-to-post hooks and confidence scoring polished for clips If those capabilities are essential, use diclip instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build me a personal substitute for diclip, not a platform clone. - Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript frontend, faster-whisper for word-level transcription, and FFmpeg for rendering. - I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the video with speaker timing. - A ranking step selects the strongest 30-90 second moments: call an LLM API key from .env with the transcript, or fall back to a local heuristic (keyword density, question marks, pauses) when no key is set. - Show each candidate clip with its hook line, a confidence score, and a one sentence reason, marked ready or pending. - Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1, or 16:9 output. - Render clips with FFmpeg using center-crop or blur-pad, burning captions styled by speaker. Never overwrite the source. Show progress and a useful failure message. - Store projects as JSON under ~/diclipDIY/projects and renders under ~/diclipDIY/exports, with a recent-projects page and a delete action. - Bind to localhost only. No accounts, telemetry, or network calls after the model and transcript are local, except the optional LLM ranking call. - Deliberately exclude face-aware reframing with seat tracking, the full in-browser timeline editor, cloud container rendering, yt-dlp downloads, and credit billing. - Add unit tests for the ranking fallback and caption grouping, plus one smoke test that imports a short fixture and produces a playable MP4. - Include a README with setup, data paths, and an honest note that CPU transcription and rendering are slow on long videos. ## Required capabilities - FFmpeg - faster-whisper - LLM API key for moment ranking (optional, falls back to heuristics) - desktop with enough storage for source media and renders ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# diclip product brief ## Problem The core loop is genuinely one-shottable: transcribe, rank the strongest moments, cut vertical clips with burned captions. The gaps are execution and infra. Face-aware reframing with seat tracking, a full in-browser timeline editor, and a container-scale render pipeline are not a one-session build. A competent agent can produce a usable personal clipper quickly, but matching diclip's reframe quality and editing depth is a multi-day project. ## Product outcome Transcribe a long video, rank the strongest moments with an LLM, and render vertical clips with burned captions using FFmpeg. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - FFmpeg - faster-whisper - LLM API key for moment ranking (optional, falls back to heuristics) - desktop with enough storage for source media and renders ## Explicit non-goals for v1 - face-aware vertical reframing with seat tracking - full in-browser timeline editor with scenes, tracks, and effects - container-scale render pipeline and yt-dlp import - credit-based capacity model and priority processing - ready-to-post hooks and confidence scoring polished for clips ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build me a personal substitute for diclip, not a platform clone. - Use Python 3.12 + FastAPI for a localhost web app, a plain JavaScript frontend, faster-whisper for word-level transcription, and FFmpeg for rendering. - I import an MP4, MOV, WebM, or MP3 file and get a transcript synced to the video with speaker timing. - A ranking step selects the strongest 30-90 second moments: call an LLM API key from .env with the transcript, or fall back to a local heuristic (keyword density, question marks, pauses) when no key is set. - Show each candidate clip with its hook line, a confidence score, and a one sentence reason, marked ready or pending. - Let me accept or reject each clip, edit in/out points, and choose 9:16, 1:1, or 16:9 output. - Render clips with FFmpeg using center-crop or blur-pad, burning captions styled by speaker. Never overwrite the source. Show progress and a useful failure message. - Store projects as JSON under ~/diclipDIY/projects and renders under ~/diclipDIY/exports, with a recent-projects page and a delete action. - Bind to localhost only. No accounts, telemetry, or network calls after the model and transcript are local, except the optional LLM ranking call. - Deliberately exclude face-aware reframing with seat tracking, the full in-browser timeline editor, cloud container rendering, yt-dlp downloads, and credit billing. - Add unit tests for the ranking fallback and caption grouping, plus one smoke test that imports a short fixture and produces a playable MP4. - Include a README with setup, data paths, and an honest note that CPU transcription and rendering are slow on long videos. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted diclip capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent
Creators pay to skip the editing grind: diclip reads the full transcript, explains why each clip was chosen, and opens a prepared edit. The reframe quality and editor depth are tuned over months, not a weekend, and the render pipeline just works on long sources.
xface-aware vertical reframing with seat tracking
xfull in-browser timeline editor with scenes, tracks, and effects
xcontainer-scale render pipeline and yt-dlp import
xcredit-based capacity model and priority processing
xready-to-post hooks and confidence scoring polished for clips
Nothing worth pointing at. That's why the prompt exists.
diclip pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0/workspace | $0/workspace | 3 clips once; 1 source-hour total; 60 min per source upload cap |
| starter | $2/workspace | — | 30 clips/mo; 2.5 source-hours/mo; 90 min per source |
| creator | $6/workspace | — | 60 clips/mo; 5.75 source-hours/mo; 120 min per source |
| studio | $11/workspace | — | 120 clips/mo; 11 source-hours/mo; 120 min per source; priority processing |
free tier3 clips once, 1 source-hour total, 60 min upload cap
billingmonthly or prepaid 3/12 months (IDR, charged in full; capacity resets monthly)
hidden costsTop-up packs (one-time, active subscription required) for extra source minutes and clips.
verified 2026-08-16 · source ↗
Is diclip free?
Free once: 3 clips and 1 source-hour total, no renewal. Paid is Creator at $6/mo (checked 2026-08-16).
Vibecode diclip
Kinda. The core of diclip is buildable in a weekend with the prompt on this page, but there are real gaps: face-aware vertical reframing with seat tracking, full in-browser timeline editor with scenes, tracks, and effects. Read the honest list above before committing.
How much does diclip cost?
diclip costs about $6/month (Creator, checked 2026-08-16), which is $72 per year.
What do I lose by replacing diclip?
Honestly: face-aware vertical reframing with seat tracking; full in-browser timeline editor with scenes, tracks, and effects; container-scale render pipeline and yt-dlp import; credit-based capacity model and priority processing; ready-to-post hooks and confidence scoring polished for clips. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to diclip?
No mature open-source alternative worth pointing at, which is exactly why the one-shot prompt on this page exists.