A thing I built and use every day

My repos have a night shift.
It reads the errors,
then ships the fix.

NightForge is one Claude Code plugin with ten skills in it. The bellows run on a schedule and keep the fire going, opening pull requests while I'm asleep. The anvils wait for me to swing, when a PR needs testing or the CI bill looks rude. Two of them fix my writing. It's public, it's free, and every rule in it got there because something broke first.

Just try it. Two commands:

claude plugin marketplace add jfreal/nightforge
claude plugin install nightforge@nightforge --scope user

10

skills in one plugin

4

bellows on a schedule, no human in the loop

128

commits since it started in August 2026

0 merges

it opens PRs. I'm the one who merges.

Why it exists

I had three error sweeps, one per app, each written by hand. In about two months they drifted into three different things. One filed issues and opened PRs. One only filed issues. One was a single sentence with no memory at all, so every night it rediscovered its whole backlog from scratch. Its reports grew to 24 KB of the same findings.

So I pulled the pipeline out into one place. The steps live in the skill and are the same for every app. Each tech stack gets an adapter: Netlify, Supabase, App Insights, GitHub. Each app gets a short card that names its repo, its logs, and how many fixes it's allowed per night.

New app? Write a card. New stack? Write an adapter. Fix the pipeline once, and every app gets the fix tomorrow morning. (Turns out "don't copy-paste the important thing" still applies when the thing is a prompt.)

The bellows

Bellows keep the fire hot while nobody's at the forge. These four run from scheduled tasks. They start with no memory of the last run, so everything they need to remember lives in a ledger, a Notion board, or the issue tracker.

Nightly

error-sweep

Pulls production errors from every source the app has, boils each one down to a stable signature, and drops anything already handled. Then it reads the actual code, files an issue, and spawns a fix agent in its own worktree that opens a PR. Quiet night? One line in the report.

Hourly

onboarding-sweep

A bot signs up as a brand-new customer and writes down everything that was wrong or slow. This sweep reads that board, turns each finding into a fix PR, and moves the finding along as the PR merges and deploys. The next signup says whether the fix held.

Hourly

coderabbit-sweep

CodeRabbit gives me roughly one review an hour across every repo. PRs that open while it's spent get a "limit reached" note and nothing ever retries them. This is the retry: once an hour, it picks exactly one starved PR (a priority-labelled one first, the oldest otherwise) and spends the review on it.

Weekly

docs-sweep

Finds every repo on my machine that uses sync-docs, runs its audit in a fresh worktree, and opens a draft PR wherever the docs drifted from the code. Clean repos get one line. No signup step: a repo joins by carrying the config.

The anvils

Some jobs should not run at 3am against a machine nobody is watching. Anvils don't swing themselves. These wait for me to bring the hammer.

/pr-test <n>

pr-test

Tests a PR the way a person would. Checks out the branch, starts the app, drives a real browser through the test plan, and ticks each box the moment it passes or fails. A box only gets ticked from something it watched happen. Reading the diff and nodding is review, not testing.

/codebase-cleanup

codebase-cleanup

Reads the whole repo first: security, correctness, error handling, dead code, types, tests, perf, accessibility, docs drift. Every finding needs a path:line as evidence and gets sized S, M, or L. Each small or medium one becomes its own PR. Big ones become an issue for me to decide.

/ci-cost-sweep

ci-cost-sweep

Finds where the CI minutes actually go, then cuts them without cutting coverage. The whole skill enforces one rule: measure, never assume. Plain mode reports. fix mode makes the changes on a branch and proves the savings on real runs.

/sync-docs

sync-docs

Source files carry a @doc:<key> tag, and the doc page that explains them carries the same key. When the code changes and the page doesn't, the audit catches it. fix repairs the page and never invents a detail to fill a gap.

The writing ones

Agents write a lot of words now: PR bodies, issues, reports. These two keep those words readable.

Always on

unslop

Cuts the tells that make text read like a machine wrote it, then puts some voice back. Removing the patterns is only half the job. Sterile writing gives itself away just as fast.

/simple-issue-description

simple-issue-description

Turns a rough bug report or a PR into a short issue about the problem and what should happen instead, with the implementation details taken out. If there's no real problem in there, it says so instead of making one up.

Bonus: output style

ELI10

Not a skill, but it lives in the repo. Every report comes back as what I did, did it work (with proof), and what I need from you. Plain words, jargon defined once, no git commands pasted at me. Built for end-of-day brains. Mine, specifically.

How it fits

NightForge is one part of the stack, where my projects run each other. The bellows run sweeps on three of them, and one skill came the other way.

What NightForge does for the others

What the others do for NightForge

Every rule got paid for

The skill files are long because the failures were quiet. A few favorites:

Green is not the same as healthy. One app served 404s to a live subscriber for 21 hours while every scheduled run reported success. The sweep only looked at exceptions, and a 404 isn't one. Now every run report has to say which kinds of failure it could and couldn't see.

Check the closed issues too. On the error sweep's first live run, a brand-new error was about to get filed. The tracker search found an issue already closed by a PR merged fourteen minutes after that error last fired. One extra search saved a junk issue and a fix agent chasing a fix that already shipped.

"Caching is good" is not a finding. A dependency cache that obviously saved time measured 71 seconds per run slower than no cache at all. The browser cache in the same workflow was a clear win. That's why ci-cost-sweep wants per-cache numbers before it touches anything.

Stale checkouts pass tests. A branch sitting in a week-old worktree serves the old code, passes the test plan, and looks exactly like a clean run. So pr-test finds the worktree, pulls it, and confirms the head commit before it starts a single service.

Setting it up

The two commands up top install all ten skills. Four work the second they land: unslop, simple-issue-description, codebase-cleanup, and ci-cost-sweep. The bellows (and pr-test and sync-docs) want a little setup first, a card or a config that names the repo, the logs, and how many fixes a night. The README on GitHub walks through each one. It's public domain, so take whatever's useful and bend it to your own stack.

The best part isn't the PRs. It's that the boring stuff finally has an owner. Every morning there's a short report and a few small diffs waiting, and my whole job is deciding which ones are good. That's a fun way to start a day.

Robots on the night shift. Humans on the merge button.

Related reading: why I spend half my build time proving it works, and how the DevOps wishlist got cheap enough to just build.