Free launch audit — a senior engineer reviews your repo in 48 hours.Claim yours
Gen2Prod
Production engineering

Cursor keeps breaking your codebase. Here is how to stop it

Every Cursor session brings its own idea of your architecture, so the codebase drifts one prompt at a time. How to measure the drift, undo it, and write rules that keep the next session in line.

GEGen2ProdProduction engineering

8 min read

A tall tower of mismatched wooden blocks leaning to one side, one block pulled halfway out

Cursor does not generate a stack the way Lovable or Bolt do. It edits yours, one prompt at a time, and it is very good at it. Three months in, though, most of the teams who come to us describe the same thing in almost the same words:

Every time I ask Cursor to fix one thing, it breaks something else.

That is not a model getting worse, and it is not you prompting badly. It is architecture drift, and it has a mechanical cause and a mechanical fix. This post is both.

Why it happens: every session starts from zero

A Cursor session knows what is in its context: your prompt, the files you have open or attached, and whatever the agent retrieves from the codebase while it works. It does not know that last Tuesday you decided data fetching lives in server components, or that lib/http.ts exists precisely so nobody calls fetch directly. Unless something tells it, it infers your conventions from the nearest example it happened to read — and copies it.

That works beautifully while the codebase has one way of doing each thing. The trouble starts the first time a session picks a different way: say, a useEffect fetch in a client component, because the file it was looking at was already a client component. Now there are two patterns. The next session that needs to fetch data finds one or the other, depending on what it read, and sometimes writes a third. Drift compounds, because every inconsistency becomes an example for the next prompt to copy.

"Fix one thing, break another" is what this looks like from the outside. The logic you asked about exists in three places. Cursor fixed the copy it found. The other two still have the bug, or now disagree with the one that was fixed.

Measure it before you fix it

Drift feels vague until you count it. These take ten minutes and turn "the codebase feels messy" into a list:

SignHow to check
One job done several waysSearch for each way of doing it — fetch(, axios, useQuery, server actions — and count the files using each
Duplicate helpersSearch for the obvious names: formatDate, cn, getUser, apiClient
any as a type strategyCount : any and as any
Dead code and dependenciesRun knip, which reports unused files, exports and packages
Half-finished refactorsLook for ButtonNew.tsx next to Button.tsx, v2/ folders, and TODO in code that ships
Tests that cannot failCovered below — this one needs more than a search

For a TypeScript project, the first four are one paste:

# How many ways do we fetch data?
grep -rlE "axios\." src | wc -l
grep -rlE "useQuery\(" src | wc -l
grep -rlE "\bfetch\(" src | wc -l
 
# Duplicate helpers
grep -rnE "(function|const) (formatDate|getUser|apiClient)\b" src
 
# any, as a number to watch go down
grep -rnE ":\s*any\b|as any\b" src --include='*.ts' --include='*.tsx' | wc -l
 
# Unused files, exports and dependencies
npx knip

Write the numbers down. They are the baseline, and the cleanup is done when they reach the targets you set — not when the code "feels better".

The tests that cannot fail

Cursor writes tests when asked, and when the real dependency is awkward it mocks it. Sometimes it mocks the very thing the test is supposed to be testing:

vi.mock("@/lib/pricing", () => ({ calculateTotal: () => 4200 }));
import { calculateTotal } from "@/lib/pricing";
 
test("calculates the cart total", () => {
  expect(calculateTotal(cart)).toBe(4200); // asserts the mock, not the code
});

That test is green today and will be green after any bug you could introduce in calculateTotal. The fast way to find these is to break things on purpose: open a function the suite claims to cover, make it return something wrong, and run the tests. If nothing goes red, the coverage is decorative. Do this for the code that earns money — pricing, checkout, permissions — before anything else.

Step 1: pin the behaviour before you touch anything

A cleanup changes the shape of the code without changing what it does, and you can only claim the second half if something checks it. Before migrating anything, put a small number of real tests on the paths that matter: an end-to-end test that signs up, does the core thing and pays, and integration tests on the server functions behind them, with nothing mocked except the outside world (email, payments, third-party APIs).

This is also the moment to delete the hollow tests rather than "fix" them. A test that asserts its own mock is not a weak test; it is a false statement about your coverage.

Step 2: pick one way per job, and write it down

For each job the codebase does more than one way, choose a winner. Usually the cheapest choice is the pattern most of the code already uses, unless that pattern is the insecure one.

JobTypical choice in a Next.js app
Reading dataServer components calling functions in lib/data/
Writing dataServer actions, each validating its input and checking the session
ValidationOne schema per input, shared by form and server
ErrorsThrown on the server, caught at one boundary per route
Types for rowsGenerated from the database schema, never hand-written
StylingWhatever already dominates — the point is one

The table is not the deliverable. The decisions are. Each row will become a line in a rules file in step 4, so phrase it the way you would explain it to a new hire: what to do, where the canonical example lives, and what not to do.

Step 3: migrate in small pull requests — Cursor can do this part

The irony of drift is that Cursor is an excellent tool for undoing it, once it knows the target. With the convention written down, the migration is repetitive, well-specified work, which is exactly what an agent is good at:

Migrate src/app/projects/ to the data-access convention in
.cursor/rules/data-access.mdc. Use src/app/invoices/page.tsx as the model.
Do not change behaviour. Do not touch files outside src/app/projects/.
When done, list every file you changed and why.

One module per pull request, one concern per pull request, reviewed by a human who reads the diff rather than skimming it. The tests from step 1 run on every one. When the count for the old pattern reaches zero, delete the old helper — and then make it impossible to bring back, which is step 5.

Step 4: the rules files

Project rules are how you tell every future session what you decided, without anyone having to remember to paste it into a prompt. They live in .cursor/rules/ as .mdc files — Markdown with a short frontmatter block — and are committed with the code, so the whole team and every session get the same instructions.

The frontmatter decides when a rule is loaded:

ModeFrontmatterLoaded when
AlwaysalwaysApply: trueEvery session, no matter what
Specific filesglobs: and alwaysApply: falseA file matching the pattern is in play
Intelligentlydescription: onlyThe agent judges the description relevant
Manuallynone of the aboveYou @-mention the rule in chat

Here is a small set that covers the decisions from step 2. First, the one that always loads — short, because it is in every context window:

---
description: Project architecture and conventions
alwaysApply: true
---
 
- Next.js App Router, TypeScript strict. No `any`; use `unknown` and narrow.
- Reading data: server components call functions in `src/lib/data/`.
  Never fetch in `useEffect`. Model: `src/app/invoices/page.tsx`.
- Writing data: server actions in `src/actions/`, built with `authedAction()`
  from `src/lib/action.ts`, which checks the session and validates input.
- Database row types come from `src/lib/db/types.ts` (generated). Never
  hand-write a row type.
- Before adding a helper, search `src/lib/` — it probably exists.
- Do not add a dependency without saying so and why.

Then rules that load only where they apply:

---
description: How tests are written in this project
globs: src/**/*.test.ts
alwaysApply: false
---
 
- Never mock the module under test. Mock only the outside world:
  email, payments, third-party HTTP. Model: `src/lib/pricing.test.ts`.
- Every test must fail if the behaviour it names is removed. If you cannot
  make it fail, say so instead of writing it.
- Do not change an assertion to make a failing test pass. Report the failure.
---
description: Reading and writing data
globs: src/lib/data/**/*.ts
alwaysApply: false
---
 
- One file per table or aggregate: `src/lib/data/projects.ts`, and so on.
- Every query that returns user data filters by the current user's id or
  organisation, inside the query — not afterwards in JavaScript.
- Return typed rows from `src/lib/db/types.ts`; never return `any`.

What separates rules that work from rules that are ignored:

  • Specific beats virtuous. "Write clean, maintainable code" changes nothing; every model already believes it is doing that. "Never fetch in useEffect" changes the next diff.
  • Point at a real file. "Model: src/app/invoices/page.tsx" gives the agent an example to copy, which is what it was going to do anyway — now it copies the right one.
  • Say what not to do. Most drift is a reasonable choice that happens to differ from yours. Name the alternatives you rejected.
  • Keep them short and split them. Cursor's own guidance is to stay under 500 lines per rule; in practice, a rule worth loading is a screenful. An always-on rule competes with your actual prompt for attention.
  • Update them in the same pull request as the convention. A rule that describes last month's architecture is worse than none, because it is confidently wrong.

Step 5: make CI enforce what the rules ask for

A rule is advice. The model reads it, usually follows it, and occasionally does not — especially deep into a long session. So everything in the rules that a machine can check should also be checked by a machine, on every pull request:

import { defineConfig } from "eslint/config";
import tseslint from "typescript-eslint";
 
export default defineConfig([
  tseslint.configs.recommended,
  {
    rules: {
      "@typescript-eslint/no-explicit-any": "error",
      "no-restricted-imports": ["error", {
        paths: [{ name: "axios", message: "Use src/lib/data/ — see .cursor/rules/architecture.mdc" }],
      }],
    },
  },
]);
name: ci
on: pull_request
jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npx tsc --noEmit
      - run: npx eslint .
      - run: npx knip
      - run: npm test

The division of labour is the whole idea. Rules make Cursor right more often. Lint, types and tests catch it when it is not, before the change reaches main. The no-restricted-imports message even points the agent back at the rule, so when it trips the check it gets told where to look.

What rules will not fix

Rules shape the next prompt. They do nothing about what is already there, and two problems in particular are invisible to a drift count:

  • Authorisation that lives in the UI. An admin page that checks user.role before rendering, calling an API route that checks nothing. The page looks protected; the route is open to anyone who finds it. It is the same failure as v0's public Server Actions, and the fix is the same: every decision on the server, with a test.
  • Secrets in the history. An .env committed in week one stays in Git after it is deleted from the tree. Removing the file is not enough; the key has to be rotated. Here is how to find the ones already leaked.

Both need someone to go looking, not a rule telling the agent to behave in future.

The order to do it in

  1. Measure: the counts above, written down as a baseline.
  2. Pin behaviour on the paths that earn money, and delete the tests that cannot fail.
  3. Decide one way per job.
  4. Migrate module by module, with Cursor doing the typing and a human reading the diffs.
  5. Write the decisions into .cursor/rules/, and enforce what you can in CI.

None of this asks you to stop using Cursor. It asks you to give it what every new developer on the team gets on day one: the conventions, written down, with an example to copy and a build that fails when they are broken.

If you would rather have the measuring done and the list handed to you, the free launch audit is a senior engineer reading your repository and coming back within 48 hours with every finding reproduced and a fixed price against each fix. What we fix in Cursor codebases covers what else we usually find, and refactoring and cleanup is the longer version of steps 2 to 5.

Share

Your AI-built product deserves a real launch.

In 48 hours you'll know what's broken and what each fix costs — before you spend a dollar.

  • Free audit, no card
  • Fixed price before we start
  • Your code stays yours