Vibecode BrandGEO
track this build5 steps, step by step0%Running a fixed battery of brand questions through five model APIs and scoring the answers with a second LLM pass is a genuine weekend build, and for one brand it answers the headline question: what does AI say about us. The gaps are the ones every tracker in this category shares. API answers approximate but do not equal the consumer apps, a score with no trend history behind it is a screenshot rather than a signal, and a rubric only becomes comparable after it has scored many brands.
You are building a lean indie version of BrandGEO. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # BrandGEO indie build ## Goal Build the smallest trustworthy replacement for the core BrandGEO workflow for one developer or a tiny team. ## Scope Send a fixed set of brand questions through the OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs, score each answer against a six-dimension rubric with a second LLM pass, store every run, render a Markdown or PDF report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - white-label PDF reports an agency can hand to a client - weekly monitoring that keeps running when nobody is thinking about it - a rubric calibrated across many brands, so scores are comparable - competitor benchmarks per brand - the consumer app surfaces, which no API exactly reproduces If those capabilities are essential, use Elmo instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local AI brand visibility auditor for one brand. Requirements: - Node 22, TypeScript, better-sqlite3, a CLI. Local only, no accounts, no telemetry. - brand.json holds my brand name, aliases, domain, one competitor, and up to 25 audit questions (who is X, best tools for Y, X vs competitor, is X legit). - `audit run` sends every question through OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs. Keys from .env; skip engines whose key is missing and say so in the report instead of failing. - Store one immutable row per run, question, and engine: raw answer, model id, latency, error text. Never overwrite a previous run. - Cap concurrency at 2 per engine and retry twice on 429 and 5xx with backoff. - A separate scoring pass grades each stored answer 0-10 on six dimensions: recognition, knowledge depth, competitive context, sentiment, contextual recall, discoverability. Rubric text lives in rubric.md, scores must cite the answer sentence that justifies them. - `audit report` renders a Markdown report: overall score per engine, the six-dimension table, competitor mentions, and every flat-out wrong claim the models made about the brand, quoted. - `audit history` prints score per engine across runs from SQLite, so week two starts meaning something. - Out of scope: white-label PDFs, multi-brand management, scheduled monitoring, and scraping the consumer web UIs. One brand, run by hand. - README: setup, per-run cost estimate by engine, and a plain note that API answers only approximate what the apps actually show users. ## Required capabilities - API keys for OpenAI, Anthropic, Gemini, xAI (Grok), and DeepSeek - a written scoring rubric with per-dimension definitions - durable per-run storage for trend history - a scheduler for recurring audits - an API budget that scales with prompts x engines ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a lean indie version of BrandGEO. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== README.md ===== # BrandGEO indie build ## Goal Build the smallest trustworthy replacement for the core BrandGEO workflow for one developer or a tiny team. ## Scope Send a fixed set of brand questions through the OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs, score each answer against a six-dimension rubric with a second LLM pass, store every run, render a Markdown or PDF report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - white-label PDF reports an agency can hand to a client - weekly monitoring that keeps running when nobody is thinking about it - a rubric calibrated across many brands, so scores are comparable - competitor benchmarks per brand - the consumer app surfaces, which no API exactly reproduces If those capabilities are essential, use Elmo instead of pretending the gap is solved. ===== AGENTS.md ===== # Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs". ===== BUILD_PLAN.md ===== # Build plan ## Original build brief Build me a local AI brand visibility auditor for one brand. Requirements: - Node 22, TypeScript, better-sqlite3, a CLI. Local only, no accounts, no telemetry. - brand.json holds my brand name, aliases, domain, one competitor, and up to 25 audit questions (who is X, best tools for Y, X vs competitor, is X legit). - `audit run` sends every question through OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs. Keys from .env; skip engines whose key is missing and say so in the report instead of failing. - Store one immutable row per run, question, and engine: raw answer, model id, latency, error text. Never overwrite a previous run. - Cap concurrency at 2 per engine and retry twice on 429 and 5xx with backoff. - A separate scoring pass grades each stored answer 0-10 on six dimensions: recognition, knowledge depth, competitive context, sentiment, contextual recall, discoverability. Rubric text lives in rubric.md, scores must cite the answer sentence that justifies them. - `audit report` renders a Markdown report: overall score per engine, the six-dimension table, competitor mentions, and every flat-out wrong claim the models made about the brand, quoted. - `audit history` prints score per engine across runs from SQLite, so week two starts meaning something. - Out of scope: white-label PDFs, multi-brand management, scheduled monitoring, and scraping the consumer web UIs. One brand, run by hand. - README: setup, per-run cost estimate by engine, and a plain note that API answers only approximate what the apps actually show users. ## Required capabilities - API keys for OpenAI, Anthropic, Gemini, xAI (Grok), and DeepSeek - a written scoring rubric with per-dimension definitions - durable per-run storage for trend history - a scheduler for recurring audits - an API budget that scales with prompts x engines ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden. ===== .env.example ===== # Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
You are building a production product version of BrandGEO. Create the following project files first, then implement the application by following them. Keep the files updated as decisions change. Do not collapse this into a single README or prompt. ===== PRODUCT.md ===== # BrandGEO product brief ## Problem Running a fixed battery of brand questions through five model APIs and scoring the answers with a second LLM pass is a genuine weekend build, and for one brand it answers the headline question: what does AI say about us. The gaps are the ones every tracker in this category shares. API answers approximate but do not equal the consumer apps, a score with no trend history behind it is a screenshot rather than a signal, and a rubric only becomes comparable after it has scored many brands. ## Product outcome Send a fixed set of brand questions through the OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs, score each answer against a six-dimension rubric with a second LLM pass, store every run, render a Markdown or PDF report. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - API keys for OpenAI, Anthropic, Gemini, xAI (Grok), and DeepSeek - a written scoring rubric with per-dimension definitions - durable per-run storage for trend history - a scheduler for recurring audits - an API budget that scales with prompts x engines ## Explicit non-goals for v1 - white-label PDF reports an agency can hand to a client - weekly monitoring that keeps running when nobody is thinking about it - a rubric calibrated across many brands, so scores are comparable - competitor benchmarks per brand - the consumer app surfaces, which no API exactly reproduces ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees. ===== ARCHITECTURE.md ===== # Architecture ## Starting brief Build me a local AI brand visibility auditor for one brand. Requirements: - Node 22, TypeScript, better-sqlite3, a CLI. Local only, no accounts, no telemetry. - brand.json holds my brand name, aliases, domain, one competitor, and up to 25 audit questions (who is X, best tools for Y, X vs competitor, is X legit). - `audit run` sends every question through OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs. Keys from .env; skip engines whose key is missing and say so in the report instead of failing. - Store one immutable row per run, question, and engine: raw answer, model id, latency, error text. Never overwrite a previous run. - Cap concurrency at 2 per engine and retry twice on 429 and 5xx with backoff. - A separate scoring pass grades each stored answer 0-10 on six dimensions: recognition, knowledge depth, competitive context, sentiment, contextual recall, discoverability. Rubric text lives in rubric.md, scores must cite the answer sentence that justifies them. - `audit report` renders a Markdown report: overall score per engine, the six-dimension table, competitor mentions, and every flat-out wrong claim the models made about the brand, quoted. - `audit history` prints score per engine across runs from SQLite, so week two starts meaning something. - Out of scope: white-label PDFs, multi-brand management, scheduled monitoring, and scraping the consumer web UIs. One brand, run by hand. - README: setup, per-run cost estimate by engine, and a plain note that API answers only approximate what the apps actually show users. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it. ===== AGENTS.md ===== # Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone. ===== MILESTONES.md ===== # Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations. ===== OPERATIONS.md ===== # Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted BrandGEO capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
# BrandGEO indie build ## Goal Build the smallest trustworthy replacement for the core BrandGEO workflow for one developer or a tiny team. ## Scope Send a fixed set of brand questions through the OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs, score each answer against a six-dimension rubric with a second LLM pass, store every run, render a Markdown or PDF report. ## Quick start 1. Install the documented dependencies. 2. Copy `.env.example` to `.env`. 3. Run the development command chosen during implementation. 4. Complete the acceptance checks in `BUILD_PLAN.md`. ## Honest limits This build deliberately does not replace: - white-label PDF reports an agency can hand to a client - weekly monitoring that keeps running when nobody is thinking about it - a rubric calibrated across many brands, so scores are comparable - competitor benchmarks per brand - the consumer app surfaces, which no API exactly reproduces If those capabilities are essential, use Elmo instead of pretending the gap is solved.
# Agent instructions - Optimize for a working, understandable weekend build. - Prefer the fewest moving parts that satisfy the brief. - Do not invent cryptography, security guarantees, APIs, or compliance claims. - Keep secrets out of source control and logs. - Add focused tests for destructive, security-sensitive, and data-loss paths. - Run the project checks before declaring the build complete. - Record any deliberate shortcut in the README under "Tradeoffs".
# Build plan ## Original build brief Build me a local AI brand visibility auditor for one brand. Requirements: - Node 22, TypeScript, better-sqlite3, a CLI. Local only, no accounts, no telemetry. - brand.json holds my brand name, aliases, domain, one competitor, and up to 25 audit questions (who is X, best tools for Y, X vs competitor, is X legit). - `audit run` sends every question through OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs. Keys from .env; skip engines whose key is missing and say so in the report instead of failing. - Store one immutable row per run, question, and engine: raw answer, model id, latency, error text. Never overwrite a previous run. - Cap concurrency at 2 per engine and retry twice on 429 and 5xx with backoff. - A separate scoring pass grades each stored answer 0-10 on six dimensions: recognition, knowledge depth, competitive context, sentiment, contextual recall, discoverability. Rubric text lives in rubric.md, scores must cite the answer sentence that justifies them. - `audit report` renders a Markdown report: overall score per engine, the six-dimension table, competitor mentions, and every flat-out wrong claim the models made about the brand, quoted. - `audit history` prints score per engine across runs from SQLite, so week two starts meaning something. - Out of scope: white-label PDFs, multi-brand management, scheduled monitoring, and scraping the consumer web UIs. One brand, run by hand. - README: setup, per-run cost estimate by engine, and a plain note that API answers only approximate what the apps actually show users. ## Required capabilities - API keys for OpenAI, Anthropic, Gemini, xAI (Grok), and DeepSeek - a written scoring rubric with per-dimension definitions - durable per-run storage for trend history - a scheduler for recurring audits - an API budget that scales with prompts x engines ## Delivery order 1. Scaffold the smallest runnable application and document its commands. 2. Implement the primary data model and core workflow. 3. Add validation, safe failure states, and persistence. 4. Cover the critical path with automated tests. 5. Exercise a clean install from the README and fix every missing step. ## Done when - A new user can go from clone to first successful workflow using only the README. - The core workflow works without paid infrastructure unless the brief requires it. - Tests cover the highest-risk behavior. - Known limitations are explicit rather than hidden.
# Copy to .env and document every variable when it is introduced. # Never put real credentials in this file. APP_ENV=development # Add only values required by the selected implementation.
# BrandGEO product brief ## Problem Running a fixed battery of brand questions through five model APIs and scoring the answers with a second LLM pass is a genuine weekend build, and for one brand it answers the headline question: what does AI say about us. The gaps are the ones every tracker in this category shares. API answers approximate but do not equal the consumer apps, a score with no trend history behind it is a screenshot rather than a signal, and a rubric only becomes comparable after it has scored many brands. ## Product outcome Send a fixed set of brand questions through the OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs, score each answer against a six-dimension rubric with a second LLM pass, store every run, render a Markdown or PDF report. ## Target user A serious builder who needs a maintainable product foundation rather than a one-off demo. ## Required capabilities - API keys for OpenAI, Anthropic, Gemini, xAI (Grok), and DeepSeek - a written scoring rubric with per-dimension definitions - durable per-run storage for trend history - a scheduler for recurring audits - an API budget that scales with prompts x engines ## Explicit non-goals for v1 - white-label PDF reports an agency can hand to a client - weekly monitoring that keeps running when nobody is thinking about it - a rubric calibrated across many brands, so scores are comparable - competitor benchmarks per brand - the consumer app surfaces, which no API exactly reproduces ## Success criteria - The primary workflow is measurable end to end. - Setup is reproducible in a clean environment. - Failure, recovery, and support paths are documented. - Product claims match what the implementation actually guarantees.
# Architecture ## Starting brief Build me a local AI brand visibility auditor for one brand. Requirements: - Node 22, TypeScript, better-sqlite3, a CLI. Local only, no accounts, no telemetry. - brand.json holds my brand name, aliases, domain, one competitor, and up to 25 audit questions (who is X, best tools for Y, X vs competitor, is X legit). - `audit run` sends every question through OpenAI, Anthropic, Gemini, xAI, and DeepSeek APIs. Keys from .env; skip engines whose key is missing and say so in the report instead of failing. - Store one immutable row per run, question, and engine: raw answer, model id, latency, error text. Never overwrite a previous run. - Cap concurrency at 2 per engine and retry twice on 429 and 5xx with backoff. - A separate scoring pass grades each stored answer 0-10 on six dimensions: recognition, knowledge depth, competitive context, sentiment, contextual recall, discoverability. Rubric text lives in rubric.md, scores must cite the answer sentence that justifies them. - `audit report` renders a Markdown report: overall score per engine, the six-dimension table, competitor mentions, and every flat-out wrong claim the models made about the brand, quoted. - `audit history` prints score per engine across runs from SQLite, so week two starts meaning something. - Out of scope: white-label PDFs, multi-brand management, scheduled monitoring, and scraping the consumer web UIs. One brand, run by hand. - README: setup, per-run cost estimate by engine, and a plain note that API answers only approximate what the apps actually show users. ## Boundaries Separate the product into replaceable modules for interface, application logic, persistence, external integrations, and operational concerns. Keep domain logic independent from delivery frameworks and vendors. ## Production baseline - Configuration: validated at startup with safe local defaults where possible. - Security: least privilege, input validation, secret redaction, rate limits on abuse-prone paths, and no invented security primitives. - Data: explicit schema and migrations, transactional writes where integrity matters, backup and restore instructions. - Integrations: adapters around third-party providers, idempotent webhook or job processing, bounded retries, and timeouts. - Observability: structured logs with request or operation IDs, an error-tracking hook, and health/readiness checks where a server exists. - Quality: unit tests for domain rules, integration tests at module boundaries, and one end-to-end critical-path test. ## Decision records For each major dependency, document why it was chosen, its failure mode, and how it can be replaced. Do not introduce infrastructure until a requirement justifies it.
# Agent instructions - Read `PRODUCT.md` and `ARCHITECTURE.md` before changing code. - Implement milestone by milestone; keep each change reviewable and leave the application runnable. - Treat authentication, payments, encryption, imports, webhooks, and destructive actions as high-risk boundaries when present. - Never invent cryptography or silently weaken a requirement to make a test pass. - Use provider interfaces for external services and deterministic fakes in tests. - Add migrations and rollback or recovery notes for persistent data changes. - Log useful operational context without credentials, tokens, passwords, or personal data. - Update documentation and run all checks before completing a milestone.
# Delivery milestones ## M0 — Decisions and scaffold - Confirm the runtime, persistence model, threat boundaries, and deployment target. - Create a reproducible local environment and continuous checks. ## M1 — Core workflow - Implement the smallest end-to-end product path with validation and tests. - Keep integrations behind interfaces. ## M2 — Trust layer - Add secure failure behavior, recovery paths, audit-relevant events, and data safeguards. - Test abuse cases and destructive operations. ## M3 — Operability - Add structured logs, error reporting hooks, health signals, backup/restore documentation, and deployment configuration. ## M4 — Release gate - Run a clean-install test, critical-path end-to-end test, dependency review, and documented rollback exercise. - Compare the shipped behavior with `PRODUCT.md` and publish remaining limitations.
# Operations ## Before release - Validate configuration and secrets at startup. - Define backup, restore, and rollback procedures and test them. - Document logs, error tracking, health signals, and alert ownership. - Set dependency update and vulnerability review expectations. ## Incident checklist 1. Contain the issue without destroying evidence or user data. 2. Record the timeline and affected scope. 3. Rotate exposed secrets and revoke compromised sessions or credentials. 4. Restore from a verified source when needed. 5. Document the root cause, remediation, and regression test. ## Launch constraint Do not market omitted BrandGEO capabilities as implemented. The v1 non-goals in `PRODUCT.md` remain user-visible limitations until they are deliberately delivered.
$ choose a build depth, inspect the files, then open the complete pack in your agent
Agencies are the tell: they pay for a report with someone else's methodology behind it that they can white-label and bill for, plus monitoring and trend history someone else keeps alive. A founder auditing one brand once is exactly who the free audit and a DIY script are for.
xwhite-label PDF reports an agency can hand to a client
xweekly monitoring that keeps running when nobody is thinking about it
xa rubric calibrated across many brands, so scores are comparable
xcompetitor benchmarks per brand
xthe consumer app surfaces, which no API exactly reproduces
Don't feel like building it? These folks already made it free.
no votes, no pay-to-list · just what's real
BrandGEO pricing
starter$79/mo · monthly · $948/yr
free tierA free audit across all five engines without a credit card, plus a 7-day trial on paid plans.
verified 2026-08-10 · source ↗
Is BrandGEO free?
A free audit across all five engines without a credit card, plus a 7-day trial on paid plans. Paid is Starter at $79/mo (checked 2026-08-10).
Vibecode BrandGEO
Kinda. The core of BrandGEO is buildable in a weekend with the prompt on this page, but there are real gaps: white-label PDF reports an agency can hand to a client, weekly monitoring that keeps running when nobody is thinking about it. Read the honest list above before committing.
How much does BrandGEO cost?
BrandGEO costs about $79/month (Starter, checked 2026-08-10), which is $948 per year.
What do I lose by replacing BrandGEO?
Honestly: white-label PDF reports an agency can hand to a client; weekly monitoring that keeps running when nobody is thinking about it; a rubric calibrated across many brands, so scores are comparable; competitor benchmarks per brand; the consumer app surfaces, which no API exactly reproduces. If any of those are load-bearing for you, keep paying.
Is there an open-source alternative to BrandGEO?
Yes: Elmo (A self-hosted AI visibility dashboard that runs your prompts across the major engines and records mentions and citations; the white-label reporting is the part you keep paying for.) The prompt is for when you want it exactly your way.