Skip to content
82 changes: 19 additions & 63 deletions AI_POLICY.md
Original file line number Diff line number Diff line change
@@ -1,73 +1,29 @@
# AI Usage Policy
# AI usage policy

The Databuddy project has strict rules for AI usage:
AI tools are welcome here. You're responsible for the work you submit: understand
it, check it, and make it useful to the people reviewing it.

- **All AI usage in any form must be disclosed.** You must state
the tool you used (e.g. Claude Code, Cursor, GitHub Copilot, ChatGPT) along with
the extent that the work was AI-assisted.
These requirements apply to outside contributions. Maintainers may use AI tools
at their discretion.

- **Pull requests created in any way by AI can only be for accepted issues.**
Drive-by pull requests that do not reference an accepted issue will be
closed. If AI isn't disclosed but a maintainer suspects its use, the
PR will be closed. If you want to share code for a non-accepted issue,
open a discussion or attach it to an existing discussion.
## Before you submit

- **Pull requests created by AI must have been fully verified with
human use.** AI must not create hypothetically correct code that
hasn't been tested. Importantly, you must not allow AI to write
code for platforms or environments you don't have access to manually
test on.
- **Disclose all AI use.** Name the tool and explain which parts of the work it helped with.
- **Start with an accepted issue.** Every AI-assisted PR must reference one. For other ideas, open a discussion or share code in an existing discussion first.
- **Review and test the complete change yourself.** AI-assisted PRs must be verified through human use. Don't submit code for a platform or environment you can't test manually.
- **Edit issues and discussions yourself.** Review and edit AI-assisted text before posting. Check the facts and remove noise.
- **Keep AI assistance to text and code.** AI-generated images, video, audio, and other media aren't accepted.

- **Issues and discussions can use AI assistance but must have a full
human-in-the-loop.** This means that any content generated with AI
must have been reviewed _and edited_ by a human before submission.
AI is very good at being overly verbose and including noise that
distracts from the main point. Humans must do their research and
trim this down.
PRs that don't meet these requirements, including undisclosed AI use, will be
closed. A maintainer will also close a PR if they suspect undisclosed AI use.
Contributors who misuse AI will be banned.

- **No AI-generated media is allowed (art, images, videos, audio, etc.).**
Text and code are the only acceptable AI-generated content, per the
other rules in this policy.
## Make review easier

- **Bad AI drivers will be banned and ridiculed in public.** You've
been warned. We love to help junior developers learn and grow, but
if you're interested in that then don't use AI, and we'll help you.
I'm sorry that bad AI drivers have ruined this for you.

These rules apply only to outside contributions to Databuddy. Maintainers
are exempt from these rules and may use AI tools at their discretion;
they've proven themselves trustworthy to apply good judgment.

## There are Humans Here

Please remember that Databuddy is maintained by humans.

Every discussion, issue, and pull request is read and reviewed by
humans (and sometimes machines, too). It is a boundary point at which
people interact with each other and the work done. It is rude and
disrespectful to approach this boundary with low-effort, unqualified
work, since it puts the burden of validation on the maintainer.

In a perfect world, AI would produce high-quality, accurate work
every time. But today, that reality depends on the driver of the AI.
And today, most drivers of AI are just not good enough. So, until either
the people get better, the AI gets better, or both, we have to have
strict rules to protect maintainers.

## AI is Welcome Here

Databuddy is written with plenty of AI assistance, and many maintainers embrace
AI tools as a productive tool in their workflow. As a project, we welcome
AI as a tool!

**Our reason for the strict AI policy is not due to an anti-AI stance**, but
instead due to the number of highly unqualified people using AI. It's the
people, not the tools, that are the problem.

I include this section to be transparent about the project's usage about
AI for people who may disagree with it, and to address the misconception
that this policy is anti-AI in nature.
Share what you tried, what you checked, and what still needs help. If you're
learning, ask questions and we'll help you find a starting point.

---

_This policy is adapted from the [Ghostty project's AI Usage Policy](https://github.com/ghostty-org/ghostty/blob/main/AI_POLICY.md), originally created by [mitchellh](https://github.com/mitchellh) and the Ghostty team. Thanks for setting the standard._
Adapted from the [Ghostty AI Usage Policy](https://github.com/ghostty-org/ghostty/blob/main/AI_POLICY.md)
by [mitchellh](https://github.com/mitchellh) and the Ghostty team.
195 changes: 52 additions & 143 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -1,178 +1,87 @@
# Contributing Guide
# Contributing to Databuddy

Thank you for your interest in contributing!
Thanks for helping make Databuddy better. Bug reports, docs fixes, and code
contributions are all welcome. For a new feature, open an issue first so we can
agree on the scope before you start.

## AI Usage
Want to run your own instance? Follow the [self-hosting guide](README.md#self-hosting).
This guide is for working on Databuddy itself.

The Databuddy project has strict rules for AI usage. Please see
the [AI Usage Policy](AI_POLICY.md). **This is very important.**
## Run locally

## 🚀 Getting Started

### Installation

1. Clone the repository:

```bash
git clone https://github.com/databuddy-analytics/Databuddy.git
cd databuddy
```

2. Install dependencies (requires Bun 1.2.0+, check with `bun --version`):

```bash
bun install
```

> [!NOTE]
> Bun 1.2.0 is the minimum for `catalog:` dependency support. The repo pins an exact version in the `packageManager` field of `package.json`; match it with `bun upgrade` or by installing that version from https://bun.sh.

3. Set up environment variables:
You'll need Docker Compose, [Node.js LTS](https://nodejs.org/en/about/previous-releases), and the Bun version pinned in
[`package.json`](package.json).

```bash
git clone --branch staging https://github.com/databuddy-analytics/Databuddy.git
cd Databuddy
bun install --frozen-lockfile
cp .env.example .env
```

4. Start Docker services (PostgreSQL, Redis, ClickHouse):
In `.env`, set `SELFHOST=true` to work without hosted billing. Set
`BETTER_AUTH_SECRET` and `DATABUDDY_ENCRYPTION_KEY` to separate random values.
The example database URLs match the local Docker services.

```bash
docker compose up -d
```

> **Note:** This starts the **development** infrastructure only (`docker-compose.yaml`).
> For self-hosting with all application services, use `docker compose -f docker-compose.selfhost.yml up -d` instead — see the [Self-Hosting section](README.md#-self-hosting) in the README.

5. Set up the database:

```bash
bun run db:push # Apply database schema
bun run clickhouse:init # Initialize ClickHouse basket
```

6. Build the SDK:

```bash
bun run sdk:build
```

7. Start development servers:
Once the databases are ready, create the schemas and start the app:

```bash
bun run db:push
bun run clickhouse:init
bun run dev:dashboard
```

8. Seed the database with sample data (optional):

```bash
bun run db:seed <WEBSITE_ID> [EVENT_COUNT]
```

**Examples:**

```bash
bun run db:seed g0zlgMtBaXzIP1EGY2ieG 10000
bun run db:seed d7zlgMtBaSzIL1EGR2ieR 5000
```

**Note:**
- The domain is automatically fetched from the database based on the website ID
- Default event count is 10,000 if not specified
- Seeds events, outgoing links, errors, and web vitals data
- You can find your website ID in your website overview settings

## 💻 Development

### Available Scripts

Check the root `package.json` for available scripts. Here are some common ones:

- `bun run dev` - Start all applications in development mode
- `bun run build` - Build all applications
- `bun run start` - Start all applications in production mode
- `bun run lint` - Lint all code with Ultracite
- `bun run format` - Format all code with Prettier
- `bun run check-types` - Type check all TypeScript code
- `bun run db:studio` - Open Drizzle Studio for database management
- `bun run db:push` - Apply database schema changes
- `bun run db:migrate` - Run database migrations
- `bun run db:deploy` - Deploy database migrations
- `bun run sdk:build` - Build the SDK package
- `bun run email:dev` - Start the email development server
Open [localhost:3000](http://localhost:3000). The dev command builds the SDK and
devtools for you. Add provider keys only for the features you're working on;
see [optional services](README.md#optional-services).

You can also `cd` into any package and run its scripts directly.
To explore with sample data, create a website, copy its ID from its settings,
and run `bun run db:seed YOUR_WEBSITE_ID 1000`.

### Development Workflow
## Check your changes

#### Branch and PR lifecycle

Keep each branch short-lived: one branch, one independently reviewable and
revertible slice, one pull request. Do not use a branch as a general work
queue.

1. Before starting, check open pull requests for the same surface, public
contract, schema, or deployment configuration. Update `staging`, then create
a fresh branch from it:

```bash
git switch staging
git pull --ff-only origin staging
git switch -c codex/short-slice
```

Do not branch from another feature branch. An exception needs an explicit
`Depends on #…` in both PRs and agreement from its owner; land the
prerequisite first.

2. Keep the branch to its stated slice. If a change can be reviewed or reverted
independently, open a separate branch and PR; leave unrelated cleanup and
refactors out of the current one.

3. Push early and open a draft PR against `staging`. State the problem being
solved and any dependency or known overlap. This makes ownership visible
before parallel work drifts into the same files.

4. Before requesting review, rebase onto the current `origin/staging` and
resolve the conflicts in the slice. Do not merge `staging` into a feature
branch just to refresh it. If the rebase changes reviewed code, request a
fresh review.

5. Run the relevant checks:
Run these from the repo root before pushing:

```bash
bun run lint
bun run check-types
bun run test
```

6. Create a changeset when the change affects a published package:

```bash
bun run changeset
```

7. Commit your changes:
Use `bun run format` for formatting with Ultracite/Biome. Package scripts in
[`package.json`](package.json) and each app's `package.json` cover other tasks.
For changes to published packages, add a changeset with `bun run changeset`.

```bash
git add .
git commit -m "feat: your feature"
```
To check self-host initialization against disposable databases, run
`bash scripts/test-selfhost-init.sh`. It needs Docker Compose 2.24.4 or later and
creates and removes its own test databases.
Comment thread
izadoesdev marked this conversation as resolved.

8. Push your changes:
## Open a pull request

```bash
git push -u origin codex/short-slice
```
1. Check open PRs for overlapping work. Start one focused branch from current `staging`:

9. When the PR is merged or closed, retire the branch. GitHub automatically
deletes merged source branches; delete a closed branch manually. Never
repurpose or reopen an old branch for a new slice—start again from current
`staging`.
```bash
git switch staging
git pull --ff-only origin staging
git switch -c codex/short-description
```

For parallel work, use one worktree per branch and never have two people or
agents mutate the same branch. Remove a worktree only after its work is merged,
closed, or safely moved to a new branch.
2. Keep each branch and PR to one change that can be reviewed and reverted on its own. Use a separate worktree for parallel work; one person or agent should edit a branch at a time.
3. Commit by intent, with a scope: `fix(api): handle missing session`. Keep unrelated changes in separate commits and PRs.
4. Push early and open a draft against `staging`. Describe the problem, what changed, how you checked it, and any dependencies or overlaps.
5. Rebase onto current `origin/staging` before review. If that changes reviewed code, request fresh review. A feature-branch dependency needs the owner's agreement and `Depends on #…` in both PRs; merge the prerequisite first.
6. Mark the PR ready when checks pass. Wait for configured reviewers on the final commit, address every actionable comment, and resolve the review threads. Before merging, check again for pending reviews or unaddressed feedback; green CI alone isn't review approval.
7. After merge or closure, retire the branch and remove its finished worktree. Merged branches are deleted automatically; delete closed branches manually. Start fresh for the next change.

Keep code simple and type-safe, and use shared components and helpers where they
fit. [AGENTS.md](AGENTS.md) has the repository conventions and full workflow.

## Code Style
## Using AI tools

- Use Biome for linting and formatting
- Follow the coding standards in the README
- Keep it simple and type-safe
AI assistance is welcome. Disclose the tool and how you used it, and personally
review and test the result. AI-assisted PRs must reference an accepted issue.
Read [AI_POLICY.md](AI_POLICY.md) before contributing with AI.
Loading