Vibecode 100 Questions
track this build5 steps, step by step0%A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
You are building a lean indie version of 100 Questions. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # 100 Questions indie build ## Goal Build the smallest trustworthy replacement for the core 100 Questions workflow for one developer or a tiny team. ## Scope Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - reliable orchestration and retries across four providers - normalized citations and evidence-linked metrics - competitor and missed-question extraction - stored point-in-time reports and comparisons - polished exports and action recommendations If those capabilities are essential, use Elmo instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local AI visibility benchmark for one brand. Requirements: - Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI. - `benchmark --domain example.com --description "..."` creates one immutable run. - Generate 25 buyer questions from the domain and description, or accept a JSON question file. - Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI. - Use each provider's supported web-search or grounding tool; keys live only in `.env`. - Limit concurrency per provider, retry transient failures, and preserve failed cells in the report. - Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite. - Detect exact and case-insensitive brand mentions; allow aliases in a config file. - Extract named competitors with one structured LLM pass after all answers are stored. - Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters. - Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources. - Show missed questions where competitors appear but the target brand does not. - Render a self-contained static HTML report with filters and expandable raw evidence. - Export questions, answer metrics, competitors, and citations as CSV files. - Every aggregate metric must link back to the answer rows used to calculate it. - Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation. - Include fixture-based tests for mention detection, URL normalization, and metric calculations. - README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. ## Required capabilities - OpenAI, Anthropic, Google Gemini, and xAI API keys - provider-specific web search or grounding tools - durable run storage - URL and citation normalization - report generation ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of 100 Questions. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # 100 Questions indie build ## Goal Build the smallest trustworthy replacement for the core 100 Questions workflow for one developer or a tiny team. ## Scope Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - reliable orchestration and retries across four providers - normalized citations and evidence-linked metrics - competitor and missed-question extraction - stored point-in-time reports and comparisons - polished exports and action recommendations If those capabilities are essential, use Elmo instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local AI visibility benchmark for one brand. Requirements: - Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI. - `benchmark --domain example.com --description "..."` creates one immutable run. - Generate 25 buyer questions from the domain and description, or accept a JSON question file. - Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI. - Use each provider's supported web-search or grounding tool; keys live only in `.env`. - Limit concurrency per provider, retry transient failures, and preserve failed cells in the report. - Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite. - Detect exact and case-insensitive brand mentions; allow aliases in a config file. - Extract named competitors with one structured LLM pass after all answers are stored. - Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters. - Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources. - Show missed questions where competitors appear but the target brand does not. - Render a self-contained static HTML report with filters and expandable raw evidence. - Export questions, answer metrics, competitors, and citations as CSV files. - Every aggregate metric must link back to the answer rows used to calculate it. - Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation. - Include fixture-based tests for mention detection, URL normalization, and metric calculations. - README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. ## Required capabilities - OpenAI, Anthropic, Google Gemini, and xAI API keys - provider-specific web search or grounding tools - durable run storage - URL and citation normalization - report generation ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of 100 Questions. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # 100 Questions product brief ## Problem A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work. ## Product outcome Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - OpenAI, Anthropic, Google Gemini, and xAI API keys - provider-specific web search or grounding tools - durable run storage - URL and citation normalization - report generation ## Explicit non-goals for v1 - reliable orchestration and retries across four providers - normalized citations and evidence-linked metrics - competitor and missed-question extraction - stored point-in-time reports and comparisons - polished exports and action recommendations ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build me a local AI visibility benchmark for one brand. Requirements: - Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI. - `benchmark --domain example.com --description "..."` creates one immutable run. - Generate 25 buyer questions from the domain and description, or accept a JSON question file. - Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI. - Use each provider's supported web-search or grounding tool; keys live only in `.env`. - Limit concurrency per provider, retry transient failures, and preserve failed cells in the report. - Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite. - Detect exact and case-insensitive brand mentions; allow aliases in a config file. - Extract named competitors with one structured LLM pass after all answers are stored. - Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters. - Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources. - Show missed questions where competitors appear but the target brand does not. - Render a self-contained static HTML report with filters and expandable raw evidence. - Export questions, answer metrics, competitors, and citations as CSV files. - Every aggregate metric must link back to the answer rows used to calculate it. - Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation. - Include fixture-based tests for mention detection, URL normalization, and metric calculations. - README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted 100 Questions capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# 100 Questions indie build ## Goal Build the smallest trustworthy replacement for the core 100 Questions workflow for one developer or a tiny team. ## Scope Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - reliable orchestration and retries across four providers - normalized citations and evidence-linked metrics - competitor and missed-question extraction - stored point-in-time reports and comparisons - polished exports and action recommendations If those capabilities are essential, use Elmo instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build me a local AI visibility benchmark for one brand. Requirements: - Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI. - `benchmark --domain example.com --description "..."` creates one immutable run. - Generate 25 buyer questions from the domain and description, or accept a JSON question file. - Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI. - Use each provider's supported web-search or grounding tool; keys live only in `.env`. - Limit concurrency per provider, retry transient failures, and preserve failed cells in the report. - Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite. - Detect exact and case-insensitive brand mentions; allow aliases in a config file. - Extract named competitors with one structured LLM pass after all answers are stored. - Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters. - Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources. - Show missed questions where competitors appear but the target brand does not. - Render a self-contained static HTML report with filters and expandable raw evidence. - Export questions, answer metrics, competitors, and citations as CSV files. - Every aggregate metric must link back to the answer rows used to calculate it. - Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation. - Include fixture-based tests for mention detection, URL normalization, and metric calculations. - README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. ## Required capabilities - OpenAI, Anthropic, Google Gemini, and xAI API keys - provider-specific web search or grounding tools - durable run storage - URL and citation normalization - report generation ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# 100 Questions product brief ## Problem A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work. ## Product outcome Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - OpenAI, Anthropic, Google Gemini, and xAI API keys - provider-specific web search or grounding tools - durable run storage - URL and citation normalization - report generation ## Explicit non-goals for v1 - reliable orchestration and retries across four providers - normalized citations and evidence-linked metrics - competitor and missed-question extraction - stored point-in-time reports and comparisons - polished exports and action recommendations ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build me a local AI visibility benchmark for one brand. Requirements: - Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI. - `benchmark --domain example.com --description "..."` creates one immutable run. - Generate 25 buyer questions from the domain and description, or accept a JSON question file. - Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI. - Use each provider's supported web-search or grounding tool; keys live only in `.env`. - Limit concurrency per provider, retry transient failures, and preserve failed cells in the report. - Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite. - Detect exact and case-insensitive brand mentions; allow aliases in a config file. - Extract named competitors with one structured LLM pass after all answers are stored. - Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters. - Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources. - Show missed questions where competitors appear but the target brand does not. - Render a self-contained static HTML report with filters and expandable raw evidence. - Export questions, answer metrics, competitors, and citations as CSV files. - Every aggregate metric must link back to the answer rows used to calculate it. - Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation. - Include fixture-based tests for mention detection, URL normalization, and metric calculations. - README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted 100 Questions capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent
They pay for a repeatable, frozen benchmark with provider failures handled, citations normalized, every metric tied to evidence, and a report that is ready to act on.
xreliable orchestration and retries across four providers
xnormalized citations and evidence-linked metrics
xcompetitor and missed-question extraction
xstored point-in-time reports and comparisons
xpolished exports and action recommendations
Don't feel like building it? These folks already made it free.
no votes, no pay-to-list · just what's real
Vibecode 100 Questions
Kinda. The core of 100 Questions is buildable in a weekend with the prompt on this page, but there are real gaps: reliable orchestration and retries across four providers, normalized citations and evidence-linked metrics. Read the honest list above before committing.
How much does 100 Questions cost?
100 Questions costs about $9/month (First benchmark, checked 2026-07-31), which is $108 per year.
What do I lose by replacing 100 Questions?
Honestly: reliable orchestration and retries across four providers; normalized citations and evidence-linked metrics; competitor and missed-question extraction; stored point-in-time reports and comparisons; polished exports and action recommendations. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to 100 Questions?
Yes: Elmo (Prompt-by-prompt AI visibility with citations and competitors across the major engines; the queries still need model or scraper credentials.) The prompt is for when you want it exactly your way.