diff --git a/AI_POLICY.md b/AI_POLICY.md index d232e55dd5..62822747bf 100644 --- a/AI_POLICY.md +++ b/AI_POLICY.md @@ -1,73 +1,29 @@ -# AI Usage Policy +# AI usage policy -The Databuddy project has strict rules for AI usage: +AI tools are welcome here. You're responsible for the work you submit: understand +it, check it, and make it useful to the people reviewing it. -- **All AI usage in any form must be disclosed.** You must state - the tool you used (e.g. Claude Code, Cursor, GitHub Copilot, ChatGPT) along with - the extent that the work was AI-assisted. +These requirements apply to outside contributions. Maintainers may use AI tools +at their discretion. -- **Pull requests created in any way by AI can only be for accepted issues.** - Drive-by pull requests that do not reference an accepted issue will be - closed. If AI isn't disclosed but a maintainer suspects its use, the - PR will be closed. If you want to share code for a non-accepted issue, - open a discussion or attach it to an existing discussion. +## Before you submit -- **Pull requests created by AI must have been fully verified with - human use.** AI must not create hypothetically correct code that - hasn't been tested. Importantly, you must not allow AI to write - code for platforms or environments you don't have access to manually - test on. +- **Disclose all AI use.** Name the tool and explain which parts of the work it helped with. +- **Start with an accepted issue.** Every AI-assisted PR must reference one. For other ideas, open a discussion or share code in an existing discussion first. +- **Review and test the complete change yourself.** AI-assisted PRs must be verified through human use. Don't submit code for a platform or environment you can't test manually. +- **Edit issues and discussions yourself.** Review and edit AI-assisted text before posting. Check the facts and remove noise. +- **Keep AI assistance to text and code.** AI-generated images, video, audio, and other media aren't accepted. -- **Issues and discussions can use AI assistance but must have a full - human-in-the-loop.** This means that any content generated with AI - must have been reviewed _and edited_ by a human before submission. - AI is very good at being overly verbose and including noise that - distracts from the main point. Humans must do their research and - trim this down. +PRs that don't meet these requirements, including undisclosed AI use, will be +closed. A maintainer will also close a PR if they suspect undisclosed AI use. +Contributors who misuse AI will be banned. -- **No AI-generated media is allowed (art, images, videos, audio, etc.).** - Text and code are the only acceptable AI-generated content, per the - other rules in this policy. +## Make review easier -- **Bad AI drivers will be banned and ridiculed in public.** You've - been warned. We love to help junior developers learn and grow, but - if you're interested in that then don't use AI, and we'll help you. - I'm sorry that bad AI drivers have ruined this for you. - -These rules apply only to outside contributions to Databuddy. Maintainers -are exempt from these rules and may use AI tools at their discretion; -they've proven themselves trustworthy to apply good judgment. - -## There are Humans Here - -Please remember that Databuddy is maintained by humans. - -Every discussion, issue, and pull request is read and reviewed by -humans (and sometimes machines, too). It is a boundary point at which -people interact with each other and the work done. It is rude and -disrespectful to approach this boundary with low-effort, unqualified -work, since it puts the burden of validation on the maintainer. - -In a perfect world, AI would produce high-quality, accurate work -every time. But today, that reality depends on the driver of the AI. -And today, most drivers of AI are just not good enough. So, until either -the people get better, the AI gets better, or both, we have to have -strict rules to protect maintainers. - -## AI is Welcome Here - -Databuddy is written with plenty of AI assistance, and many maintainers embrace -AI tools as a productive tool in their workflow. As a project, we welcome -AI as a tool! - -**Our reason for the strict AI policy is not due to an anti-AI stance**, but -instead due to the number of highly unqualified people using AI. It's the -people, not the tools, that are the problem. - -I include this section to be transparent about the project's usage about -AI for people who may disagree with it, and to address the misconception -that this policy is anti-AI in nature. +Share what you tried, what you checked, and what still needs help. If you're +learning, ask questions and we'll help you find a starting point. --- -_This policy is adapted from the [Ghostty project's AI Usage Policy](https://github.com/ghostty-org/ghostty/blob/main/AI_POLICY.md), originally created by [mitchellh](https://github.com/mitchellh) and the Ghostty team. Thanks for setting the standard._ +Adapted from the [Ghostty AI Usage Policy](https://github.com/ghostty-org/ghostty/blob/main/AI_POLICY.md) +by [mitchellh](https://github.com/mitchellh) and the Ghostty team. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 43362c2881..4de55f1a7a 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,178 +1,87 @@ -# Contributing Guide +# Contributing to Databuddy -Thank you for your interest in contributing! +Thanks for helping make Databuddy better. Bug reports, docs fixes, and code +contributions are all welcome. For a new feature, open an issue first so we can +agree on the scope before you start. -## AI Usage +Want to run your own instance? Follow the [self-hosting guide](README.md#self-hosting). +This guide is for working on Databuddy itself. -The Databuddy project has strict rules for AI usage. Please see -the [AI Usage Policy](AI_POLICY.md). **This is very important.** +## Run locally -## πŸš€ Getting Started - -### Installation - -1. Clone the repository: - -```bash -git clone https://github.com/databuddy-analytics/Databuddy.git -cd databuddy -``` - -2. Install dependencies (requires Bun 1.2.0+, check with `bun --version`): - -```bash -bun install -``` - -> [!NOTE] -> Bun 1.2.0 is the minimum for `catalog:` dependency support. The repo pins an exact version in the `packageManager` field of `package.json`; match it with `bun upgrade` or by installing that version from https://bun.sh. - -3. Set up environment variables: +You'll need Docker Compose, [Node.js LTS](https://nodejs.org/en/about/previous-releases), and the Bun version pinned in +[`package.json`](package.json). ```bash +git clone --branch staging https://github.com/databuddy-analytics/Databuddy.git +cd Databuddy +bun install --frozen-lockfile cp .env.example .env ``` -4. Start Docker services (PostgreSQL, Redis, ClickHouse): +In `.env`, set `SELFHOST=true` to work without hosted billing. Set +`BETTER_AUTH_SECRET` and `DATABUDDY_ENCRYPTION_KEY` to separate random values. +The example database URLs match the local Docker services. ```bash docker compose up -d ``` -> **Note:** This starts the **development** infrastructure only (`docker-compose.yaml`). -> For self-hosting with all application services, use `docker compose -f docker-compose.selfhost.yml up -d` instead β€” see the [Self-Hosting section](README.md#-self-hosting) in the README. - -5. Set up the database: - -```bash -bun run db:push # Apply database schema -bun run clickhouse:init # Initialize ClickHouse basket -``` - -6. Build the SDK: - -```bash -bun run sdk:build -``` - -7. Start development servers: +Once the databases are ready, create the schemas and start the app: ```bash +bun run db:push +bun run clickhouse:init bun run dev:dashboard ``` -8. Seed the database with sample data (optional): - -```bash -bun run db:seed [EVENT_COUNT] -``` - -**Examples:** - -```bash -bun run db:seed g0zlgMtBaXzIP1EGY2ieG 10000 -bun run db:seed d7zlgMtBaSzIL1EGR2ieR 5000 -``` - -**Note:** -- The domain is automatically fetched from the database based on the website ID -- Default event count is 10,000 if not specified -- Seeds events, outgoing links, errors, and web vitals data -- You can find your website ID in your website overview settings - -## πŸ’» Development - -### Available Scripts - -Check the root `package.json` for available scripts. Here are some common ones: - -- `bun run dev` - Start all applications in development mode -- `bun run build` - Build all applications -- `bun run start` - Start all applications in production mode -- `bun run lint` - Lint all code with Ultracite -- `bun run format` - Format all code with Prettier -- `bun run check-types` - Type check all TypeScript code -- `bun run db:studio` - Open Drizzle Studio for database management -- `bun run db:push` - Apply database schema changes -- `bun run db:migrate` - Run database migrations -- `bun run db:deploy` - Deploy database migrations -- `bun run sdk:build` - Build the SDK package -- `bun run email:dev` - Start the email development server +Open [localhost:3000](http://localhost:3000). The dev command builds the SDK and +devtools for you. Add provider keys only for the features you're working on; +see [optional services](README.md#optional-services). -You can also `cd` into any package and run its scripts directly. +To explore with sample data, create a website, copy its ID from its settings, +and run `bun run db:seed YOUR_WEBSITE_ID 1000`. -### Development Workflow +## Check your changes -#### Branch and PR lifecycle - -Keep each branch short-lived: one branch, one independently reviewable and -revertible slice, one pull request. Do not use a branch as a general work -queue. - -1. Before starting, check open pull requests for the same surface, public - contract, schema, or deployment configuration. Update `staging`, then create - a fresh branch from it: - -```bash -git switch staging -git pull --ff-only origin staging -git switch -c codex/short-slice -``` - - Do not branch from another feature branch. An exception needs an explicit - `Depends on #…` in both PRs and agreement from its owner; land the - prerequisite first. - -2. Keep the branch to its stated slice. If a change can be reviewed or reverted - independently, open a separate branch and PR; leave unrelated cleanup and - refactors out of the current one. - -3. Push early and open a draft PR against `staging`. State the problem being - solved and any dependency or known overlap. This makes ownership visible - before parallel work drifts into the same files. - -4. Before requesting review, rebase onto the current `origin/staging` and - resolve the conflicts in the slice. Do not merge `staging` into a feature - branch just to refresh it. If the rebase changes reviewed code, request a - fresh review. - -5. Run the relevant checks: +Run these from the repo root before pushing: ```bash +bun run lint +bun run check-types bun run test ``` -6. Create a changeset when the change affects a published package: - -```bash -bun run changeset -``` - -7. Commit your changes: +Use `bun run format` for formatting with Ultracite/Biome. Package scripts in +[`package.json`](package.json) and each app's `package.json` cover other tasks. +For changes to published packages, add a changeset with `bun run changeset`. -```bash -git add . -git commit -m "feat: your feature" -``` +To check self-host initialization against disposable databases, run +`bash scripts/test-selfhost-init.sh`. It needs Docker Compose 2.24.4 or later and +creates and removes its own test databases. -8. Push your changes: +## Open a pull request -```bash -git push -u origin codex/short-slice -``` +1. Check open PRs for overlapping work. Start one focused branch from current `staging`: -9. When the PR is merged or closed, retire the branch. GitHub automatically - deletes merged source branches; delete a closed branch manually. Never - repurpose or reopen an old branch for a new sliceβ€”start again from current - `staging`. + ```bash + git switch staging + git pull --ff-only origin staging + git switch -c codex/short-description + ``` -For parallel work, use one worktree per branch and never have two people or -agents mutate the same branch. Remove a worktree only after its work is merged, -closed, or safely moved to a new branch. +2. Keep each branch and PR to one change that can be reviewed and reverted on its own. Use a separate worktree for parallel work; one person or agent should edit a branch at a time. +3. Commit by intent, with a scope: `fix(api): handle missing session`. Keep unrelated changes in separate commits and PRs. +4. Push early and open a draft against `staging`. Describe the problem, what changed, how you checked it, and any dependencies or overlaps. +5. Rebase onto current `origin/staging` before review. If that changes reviewed code, request fresh review. A feature-branch dependency needs the owner's agreement and `Depends on #…` in both PRs; merge the prerequisite first. +6. Mark the PR ready when checks pass. Wait for configured reviewers on the final commit, address every actionable comment, and resolve the review threads. Before merging, check again for pending reviews or unaddressed feedback; green CI alone isn't review approval. +7. After merge or closure, retire the branch and remove its finished worktree. Merged branches are deleted automatically; delete closed branches manually. Start fresh for the next change. +Keep code simple and type-safe, and use shared components and helpers where they +fit. [AGENTS.md](AGENTS.md) has the repository conventions and full workflow. -## Code Style +## Using AI tools -- Use Biome for linting and formatting -- Follow the coding standards in the README -- Keep it simple and type-safe +AI assistance is welcome. Disclose the tool and how you used it, and personally +review and test the result. AI-assisted PRs must reference an accepted issue. +Read [AI_POLICY.md](AI_POLICY.md) before contributing with AI. diff --git a/README.md b/README.md index 64cc0c3fc6..e2409053fb 100644 --- a/README.md +++ b/README.md @@ -1,147 +1,102 @@ # Databuddy -
+Understand how people use your product: where they come from, what they do, and +where they drop off. Use that insight to decide what to build or improve next. -[![License: AGPL](https://img.shields.io/badge/License-AGPL-red.svg)](LICENSE) -[![TypeScript](https://img.shields.io/badge/TypeScript-5.9-blue.svg)](https://www.typescriptlang.org/) -[![Next.js](https://img.shields.io/badge/Next.js-16.1-black.svg)](https://nextjs.org/) -[![React](https://img.shields.io/badge/React-19.2-blue.svg)](https://reactjs.org/) -[![Turborepo](https://img.shields.io/badge/Turborepo-2.7-blue.svg)](https://turbo.build/repo) -[![Bun](https://img.shields.io/badge/Bun-1.3-blue.svg)](https://bun.sh/) -[![Tailwind CSS](https://img.shields.io/badge/Tailwind-4.1-blue.svg)](https://tailwindcss.com/) +- **Building a product?** [Try hosted Databuddy](https://app.databuddy.cc) or follow the [tracker setup guide](https://www.databuddy.cc/docs/getting-started). +- **Running your own stack?** Start with [self-hosting](#self-hosting) below. +- **Want to help build Databuddy?** Read the [contributor guide](CONTRIBUTING.md). Bug reports and docs fixes count too. -[![CodeRabbit Pull Request Reviews](https://img.shields.io/coderabbit/prs/github/databuddy-analytics/Databuddy?utm_source=oss&utm_medium=github&utm_campaign=databuddy-analytics%2FDatabuddy&labelColor=171717&color=FF570A&link=https%3A%2F%2Fcoderabbit.ai&label=CodeRabbit+Reviews)](https://coderabbit.ai) -[![Code Coverage](https://img.shields.io/badge/coverage-85%25-green.svg)](https://github.com/databuddy-analytics/Databuddy/actions/workflows/coverage.yml) -[![Security Scan](https://img.shields.io/badge/security-A%2B-green.svg)](https://github.com/databuddy-analytics/Databuddy/actions/workflows/security.yml) -[![Dependency Status](https://img.shields.io/badge/dependencies-up%20to%20date-green.svg)](https://github.com/databuddy-analytics/Databuddy/actions/workflows/dependencies.yml) +## Self-hosting -[Vercel OSS Program](https://vercel.com/oss) - -[![Discord](https://img.shields.io/badge/Discord-Join-blue?logo=discord)](https://discord.gg/JTk7a38tCZ) -[![GitHub Stars](https://img.shields.io/github/stars/databuddy-analytics/Databuddy?style=social)](https://github.com/databuddy-analytics/Databuddy/stargazers) -[![Twitter](https://img.shields.io/twitter/follow/trydatabuddy?style=social)](https://twitter.com/trydatabuddy) - -
- -A comprehensive analytics and data management platform built with Next.js, TypeScript, and modern web technologies. Databuddy provides real-time analytics, user tracking, and data visualization capabilities for web applications. - -## 🌟 Features +Run Databuddy on your server with Docker Compose. It sets `SELFHOST=true`, so +events go straight to ClickHouse, while hosted billing and Databuddy's own telemetry are disabled. +Email and AI are optional; see [optional services](#optional-services). -- πŸ“Š Real-time analytics dashboard -- πŸ‘₯ User behavior tracking -- πŸ“ˆ Advanced data visualization // Soon -- πŸ”’ Secure authentication -- πŸ“± Responsive design -- 🌐 Multi-tenant support -- πŸ”„ Real-time updates // Soon -- πŸ“Š Custom metrics // Soon -- 🎯 Goal tracking -- πŸ“ˆ Conversion analytics -- πŸ” Custom event tracking -- πŸ“Š Funnel analysis -- πŸ“ˆ Cohort analysis // Soon -- πŸ”„ A/B testing // Soon -- πŸ“ˆ Export capabilities -- πŸ”’ GDPR compliance -- πŸ” Data encryption -- πŸ“Š API access +### Start your instance -## πŸ“š Table of Contents +These steps need a release with the `databuddy-init` image. None is published yet; +check [releases](https://github.com/databuddy-analytics/Databuddy/releases) before starting. +You'll need Git and Docker Compose; Bun and Node are only needed for +[local development](CONTRIBUTING.md#run-locally). -1. **How do I get started?** - Follow the [Getting Started](https://www.databuddy.cc/docs/getting-started) guide. -- [Contributing](#-contributing) -- [Security](#-security) -- [FAQ](#-faq) -- [Support](#-support) -- [License](#-license) - -### Prerequisites - -- Bun 1.3.14+ -- Node.js 20+ +```bash +git clone https://github.com/databuddy-analytics/Databuddy.git +cd Databuddy +git checkout YOUR_RELEASE_TAG +cp selfhost.env.example .env +``` -## 🏠 Self-Hosting +In `.env`, set: -Databuddy can be self-hosted using Docker Compose. The repo includes two compose files: +- `IMAGE_TAG` to the release you checked out. +- `POSTGRES_PASSWORD`, `CLICKHOUSE_PASSWORD`, and `REDIS_PASSWORD` to URL-safe passwords. +- `BETTER_AUTH_SECRET` and `DATABUDDY_ENCRYPTION_KEY` to separate random secrets. -| File | Purpose | -|---|---| -| `docker-compose.yaml` | **Development only** β€” starts infrastructure (Postgres, ClickHouse, Redis) for local dev | -| `docker-compose.selfhost.yml` | **Production / self-hosting** β€” backend services from GHCR images | +The template includes local URLs. Compose supplies database connections, +`SELFHOST`, and browser settings; you don't need to repeat them in `.env`. -### Quick Start +Generate each password and secret separately with `openssl rand -hex 32`. +Then start Databuddy: ```bash -# 1. Configure environment -cp .env.example .env -# Edit .env β€” set IMAGE_TAG, URL-safe database/cache passwords, public URLs, -# BETTER_AUTH_SECRET, DATABUDDY_ENCRYPTION_KEY, IP_HASH_SALT, and -# AI_GATEWAY_API_KEY. Make the local database URLs use the same credentials -# before running the initialization commands below. - -# 2. Start databases and cache -docker compose -f docker-compose.selfhost.yml up -d postgres clickhouse redis - -# 3. Initialize databases from the repo checkout (first run only) -bun install --frozen-lockfile -bun run db:push -bun run clickhouse:init - -# 4. Start backend services -docker compose -f docker-compose.selfhost.yml up -d -``` - -Services started: -- **API** β†’ `localhost:3001` -- **Basket** (event ingestion) β†’ `localhost:4000` -- **Insights** (investigation worker) β†’ `localhost:4002` -- **Links** (short links) β†’ `localhost:2500` - -All ports are configurable via env vars (`API_PORT`, `BASKET_PORT`, etc.). See the compose file comments for the full env var reference. +# Start the databases and create their schemas +docker compose -f docker-compose.selfhost.yml run --rm init -## 🀝 Contributing +# Build the dashboard for your URLs and start the apps +docker compose -f docker-compose.selfhost.yml up -d --build +``` -See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines. +Open your dashboard URL, create an account, and add your first website. +The stack includes the dashboard, API, Basket event collector, and short-link +service (port `2500`). Ports are configurable in `docker-compose.selfhost.yml`. -## πŸ”’ Security +For a public instance, replace the template's local URLs with your HTTPS URLs. +Keep the dashboard and API on the same parent domain. Set `BETTER_AUTH_COOKIE_DOMAIN`, such as `.example.com`, +to share login across subdomains. Leave it empty for localhost. Rebuild the +dashboard after changing public URLs; they're part of its browser bundle. -See [SECURITY.md](SECURITY.md) for reporting vulnerabilities. +### Optional services -## ❓ FAQ +- **Email:** For resets, invitations, and alerts, set `RESEND_API_KEY` and an `EMAIL_FROM` sender on your verified domain, such as `Databuddy `. Leave `ALERTS_EMAIL_FROM` empty to use the same sender. Recreate the services after changes. +- **Insights:** Set `AI_GATEWAY_API_KEY` and `COMPOSE_PROFILES=insights` in `.env`, then rerun `docker compose -f docker-compose.selfhost.yml up -d --build`. Website research also needs `FIRECRAWL_API_KEY`. +- **Status pages:** Deploy [the status app](apps/status) separately. Build it with `NEXT_PUBLIC_SELFHOST=true`, `NEXT_PUBLIC_API_URL` set to your API URL, and `NEXT_PUBLIC_STATUS_URL` set to its own public URL; they're part of its browser bundle. Then set `STATUS_URL` to that same URL in `.env` and rebuild the dashboard. Public status links stay hidden until you configure it. +- **DQL:** Requires separate setup: a restricted `dql_user` and `CLICKHOUSE_DQL_URL` passed to the API in Compose. Use HTTPS outside loopback and never use the application's admin credentials. See the [DQL setup script](packages/db/src/clickhouse/dql.ts). -### General +Self-hosting is still evolving. If you get stuck, [tell us what happened](https://github.com/databuddy-analytics/Databuddy/issues) or ask in [Discord](https://discord.gg/JTk7a38tCZ). -1. **What is Databuddy?** - Databuddy is a comprehensive analytics and data management platform. +### Upgrade your instance -2. **How do I get started?** - Follow the [Getting Started](https://www.databuddy.cc/docs/getting-started) guide. +Back up your databases and `.env`, check out the new release in the same directory, +and update `IMAGE_TAG`. +Keep your existing `DATABUDDY_ENCRYPTION_KEY` so stored data stays readable. +Pull the images, then apply PostgreSQL changes so you can review any prompts: -3. **Is it free?** - Check our [pricing page](https://databuddy.cc/pricing). +```bash +docker compose -f docker-compose.selfhost.yml pull --ignore-buildable +docker compose -f docker-compose.selfhost.yml pull init +docker compose -f docker-compose.selfhost.yml run --rm init bun run --cwd packages/db db:push +``` -### Technical +If you decline a change, stop the upgrade. After accepting the changes, create +any missing ClickHouse tables and views: -1. **What are the system requirements?** - See [Prerequisites](#prerequisites). +```bash +docker compose -f docker-compose.selfhost.yml run --rm init bun --cwd packages/db src/clickhouse/setup.ts +``` -2. **How do I deploy?** - See the deployment documentation in our [docs](https://databuddy.cc/docs). +This creates missing objects; it doesn't update existing ones. Apply any extra +migrations in the release notes before starting the updated apps with +`docker compose -f docker-compose.selfhost.yml up -d --build`. -3. **How do I contribute?** - See [Contributing](#contributing). +## Stay in touch -## πŸ’¬ Support +[Docs](https://www.databuddy.cc/docs) Β· [Discord](https://discord.gg/JTk7a38tCZ) Β· [GitHub issues](https://github.com/databuddy-analytics/Databuddy/issues) Β· [Email](mailto:support@databuddy.cc) -- [Documentation](https://www.databuddy.cc/docs) -- [Discord](https://discord.gg/JTk7a38tCZ) -- [Twitter](https://twitter.com/trydatabuddy) -- [GitHub Issues](https://github.com/databuddy-analytics/Databuddy/issues) -- [Email Support](mailto:support@databuddy.cc) +Found a security issue? Please follow [SECURITY.md](SECURITY.md). -## πŸ“„ License +## License -This project is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0). See the [LICENSE](LICENSE) file for details. +[AGPL-3.0](LICENSE). Copyright (c) 2025 Databuddy Analytics, Inc. -Copyright (c) 2025 Databuddy Analytics, Inc. +[Vercel OSS Program](https://vercel.com/oss)