A markdown file in the repository, rendered.

barerepo / server / SPEC-NOTES.md
rendered
log files threads runs releases config jump to file t

Book amendments

Thirteen contradictions were found reading the book against plans/docs/BUILD.md, plans/build.py and the 24 mockups. All are resolved. The book is the source of truth, so the book was changed; this file records what changed and why, and is not a second specification.

Every resolution took the side that costs the user least.

Resolved in the book

1. The security chapter reference. README.md sent readers to chapter 36, which is Proposals. Now chapter 42, which is the security chapter.

2. build.py now generates every page. Five pages were hand-written and outside the generator: signup, profile, repo-log, repo-config, thread. They are in PAGES now, so python3 build.py writes all 25 files from one shared chrome, which is what the README always claimed. Their drift went with them: the card height was wrong by 24px and the separator glyph differed.

3. Deleting a repository. 33.12 said "gone at once", 44.4 said a 30-day trash window. The window wins, because an instant irreversible delete produces a support request forge has no channel to answer. 33.12 now states the window, and states that it is a window to notice a mistake rather than a backup.

4. One name for the push limit. [limits] max_push_size_mb = 512 and [behavior] max_push_mb = 2048 were the same setting. Now [limits] max_push_mb = 2048, with a rule that removes the ambiguity for every future key: numbers live in [limits], switches live in [behavior]. 2048 wins over 512 because the push most likely to hit this limit is somebody's first import of an existing repository, and rejecting that is the worst possible first contact.

5. The pre-receive pseudocode. Appendix D nested the notes-namespace check and the catch-all rejection inside the blob-size loop, so a push to refs/notes/threads/* was never authorised and the catch-all never ran on a push that added no blobs. The ref chain is now exhaustive and size is judged once, after the refs, over the objects the push actually adds.

6. The reserved-name list. 42.5 listed eleven names but missed inbox, tokens and auth, all of which are routes an account name could shadow. The list is now every routed name plus four held for later, and the book says to derive it from the route table in code rather than copy it.

7. Raw file content has a route. 42.3 required a separate domain but no route existed and two mockups linked to one. Added GET /<user>/<repo>/raw/<ref>/<path> and [server] raw_url. When raw_url is empty, raw is served from the main host as a download with the four hardening headers, and the handler reads no session cookie, so it answers for public repositories only. Most people install on one hostname; a link that only works for operators who own a second domain is a link most users never get.

8. The server has a command. Appendix E named the ssh and hook entry points but never the daemon. forge serve now runs the server, and the ssh entry point that only authorized_keys invokes is forge ssh. The command a person types got the obvious name.

9. repo-config.html showed invalid TOML. uproar.local = [...] unquoted is a dotted key, meaning table uproar, key local. Now quoted, matching chapter 14. It matters on the one page whose entire point is that the config is a file you edit by hand.

10. The thread page shows the offline commands. Chapters 13 and 24 both require them and the mockup had none. They are the proof of the portability claim, on the page making the claim.

11. The keys page has feed tokens. Chapters 19.5 and 39.3 issue and revoke them there. The page now shows all three credentials and says plainly what each one can do, because a user pasting a token into a feed reader should know it cannot write.

12. One separator glyph. · throughout. These pages are full of diff stats where a hyphen is already a minus sign.

13. Archiving was a one-way door. Found while rewriting the hook pseudocode. 21.3 says unarchive by editing one line and pushing, but archiving rejected every push, so nothing could ever be unarchived. The owner is now exempt, for the same reason the owner can always push a broken config: owner access comes from the namespace, not from the file.

Found while building stage 1

14. An empty repository is private, and nothing else can be. Chapter 11 says push-to-create sets visibility private, but visibility is [repo] visibility in .barerepo/config, and a repository created by a push has no tree to hold that file. Appendix B's schema default was public, so a pushed-to-create repository was world-readable. A live clone confirmed the leak.

Resolved: absent means private, and only the exact word public opens a repository. An empty value, an unknown value and a typo all stay closed. The web form's visibility field stores nothing at creation either; choosing public adds the line that sets it to the block the empty-repository page tells you to paste. Chapters 11, 18, 24 and appendix B all say so now.

15. A push must be challenged before it is answered. A git client sends no credential until it gets a 401. On a public repository the ref advertisement for a push succeeded anonymously, so the client never sent its token and the push was then refused for the wrong reason: "you cannot push here" when the truth was "you were never asked who you are". Chapter 41.4 now says to return 401 on both halves of a push, and to never challenge a read.

Found while building stage 2

16. The file view links to a history page that has no route. repo-file.html and repo-config.html both show history · raw in the footer. Chapter 24 says the config page has "history, blame, and raw links". Appendix C routes neither history nor, until amendment 7, raw.

History would be the log filtered to one path, which git answers with git log -- <path> and which nothing else in the book describes. Left unbuilt and rendered as plain muted text rather than a link, so the footer keeps its shape without offering a page that is not there. It needs a route in appendix C before it can be built.

17. The compare mockup has no submit button. Chapter 24 says both sides are free text fields, and every other form mockup shows a button. repo-compare.html shows the two refs as boxes with nothing to press. Built with a compare button in the same style as the other forms, because a form a keyboard user can submit and a mouse user cannot is not finished.

Found while building stage 3

18. The mockups are built from inline styles, which the product's own policy forbids. Chapter 42.7 sets style-src 'self' with no 'unsafe-inline', so a style="..." attribute is dropped by the browser. Every mockup in plans/ is made almost entirely of them.

That is fine for the mockups: they are opened as files, with no policy, and inline style is a reasonable way to write a static reference. It is not fine for the real pages, and it fails silently, which is the dangerous part. A template that carries a style attribute renders with that spacing missing and nothing says so.

Caught by measuring a live page against its mockup: the gap under the sign-in button was 0px where the mockup has 18px, because the attribute holding it was being dropped.

Every template now uses classes only, and TestNoInlineStyles fails the build if a style attribute or a script tag reappears. Anyone porting a mockup to a template has to move its spacing into barerepo.css on the way.

19. The keys mockup shows token values it cannot possibly know. keys.html listed rt_live_7Kq2mXe and ft_live_9Xk2m as the headline of each token row. Chapter 15 says a token is displayed once at creation, inside the command that uses it, and that the server stores a hash. So the keys page can never render either string.

The mockup now names what the row actually is: the machine a runner token is attached to, or the feed a feed token opens. The token value appears exactly once, on the page that issued it, with a line saying it will not be shown again.

20. Signing in did not record which key was used. Chapter 32.5 says the keys page shows a last-used time per key, "check this if you think a key is lost". ssh-keygen -Y verify answers only yes or no, so an allowed_signers file holding every key cannot say which one matched. The keys are now tried one at a time and the matching one is stamped.

Still missing: a push over ssh does not stamp the key either, because authorized_keys passes only --account. Adding the fingerprint to that forced command would fix it, and chapter 41.3 would need to say so.

Found while building stage 4

21. Appendix A's revision refs cannot exist. The layout said:

refs/proposals/<n>              a proposal
refs/proposals/<n>/rev/<k>      a retained earlier revision

A ref is a file, so those two ask git for a file and a directory of the same name. git refuses outright:

cannot lock ref 'refs/proposals/47/rev/1':
  'refs/proposals/47' exists; cannot create 'refs/proposals/47/rev/1'

refs/proposals/47 cannot move, because chapter 36.5 tells every reviewer to fetch it by that name. So the revisions moved: refs/revisions/<n>/<k>. Appendix A, chapter 12, chapter 26 and chapter 43.5 all say so now, and a test creates both refs so the layout cannot drift back.

22. Appendix D says rewrite(ref -> ...) without saying how. A pre-receive hook cannot redirect a push; it can only accept or reject. The mechanism is git's proc-receive hook, and the config that enables it is receive.procReceiveRefs with a prefix value, not a glob.

Getting that wrong fails silently: with refs/proposals/* the hook is never called, no error appears anywhere, and the push creates a ref literally named refs/proposals/new, the one thing appendix A says never exists. Chapter 12 now documents the mechanism and the trap.

23. Chapter 13's note layout could not be read by git. It said the note tree holds one blob per comment, named <unix-timestamp>-<author>-<short-hash>.

git notes looks a note up by the hash of the object it annotates, so the tree must be keyed by object hash. A tree keyed by anything else is invisible to every command chapter 35.4 tells the user to run, and the offline promise in chapter 3 stops being true.

Checked both halves against real git rather than assumed:

  • git notes --ref=threads/47 add writes the blob at the path <sha>.
  • A meta blob alongside it does not disturb git log --show-notes=threads/47, because meta is not a valid hash and git ignores it.

Chapter 13 now says: one note per annotated object, comments as records inside it separated by --, and meta at the tree root for title, state and attached ref. Records are also the shape union merge resolves correctly, which is what chapter 7 already asked for.

24. The sign-in page told users to run a program they do not have. The mockup's only instruction was forge auth john. Signing a nonce needs nothing but OpenSSH, which is already installed, and Part VII opens by promising that no task needs the forge CLI.

Chapter 31.3 also made it harder than it is: three commands, writing the nonce to /tmp/nonce, leaving /tmp/nonce.sig behind, and using echo -n, which is not portable between shells. ssh-keygen -Y sign reads standard input and writes to standard output, so it is one line and leaves nothing:

printf '%s' '<nonce>' | ssh-keygen -Y sign -f ~/.ssh/id_ed25519 -n barerepo-auth -

Verified by copying the command the live page prints, verbatim, and signing in with it.

Rule 2 in chapter 5 now says this outright: the command a page shows must be one the user can already run, so the plain form comes first and forge is offered second as the shortcut it is. TestTheCliIsNeverTheOnlyWay fails the build if a template names a forge subcommand without its plain equivalent.

The runner page is the one honest exception, because a runner long-polls and a shell one-liner cannot. It now says so, and gives the four HTTP endpoints so anyone can write their own.

25. The diff mockups have no line numbers, but chapter 35.3 says to press one. "Open the proposal diff. Press the line number. Type your comment." The diff rows in repo-commit.html, repo-compare.html and repo-log.html render the code and nothing else, so there was no number to press, and an anchor of config.go:43 pointed at a 43 the reader could not see.

The diff helper in build.py now numbers each line on the new side. A removed line gets a blank gutter, because it has no number on that side. The real diff parser carries the same number, and the number is a link wherever there is a thread to attach a comment to.

26. Build results were unreadable by git, and the layout was a dead end. Appendix A put them at refs/notes/runs/<sha>, one notes ref per built commit. Chapter 16 claims build history clones and is readable. Both halves fail:

  • git log --show-notes=runs prints nothing, because git looks a note up by the hash of the object it annotates, not by the ref's name. Tested.
  • refs/notes/runs and refs/notes/runs/<sha> cannot both exist, so a repository that used per-commit refs can never move to the working layout without deleting every one of them first. Tested, and git says so plainly: 'refs/notes/runs/965790c...' exists; cannot create 'refs/notes/runs'.

A busy repository would also carry one ref per commit it ever built.

Now refs/notes/runs, one ref, tree keyed by the built commit. Several runs of one commit are separate records, split by --, the same shape chapter 13 uses for comments. refs/notes/releases/<tag> moved to refs/notes/releases for the same reason, keyed by the tag object.

Chapter 16 also said a log under 64kb "is inlined" without saying where. It is an output field now, and log names a blob when the output is larger.

27. A run record could not say what was built. Chapter 24 says the runs page shows "commit, status, duration, ref, and which machine ran it", and runs.html puts the ref on every row. Chapter 16's record had no ref field.

A commit arrives on a branch and on a proposal, so the commit alone does not answer it. The record carries ref now.

28. The content security policy blocked the JavaScript the book budgets for. Chapter 42.7 set default-src 'none' with no script-src, which blocks every script, and closed with "if a future feature needs script-src, that feature is wrong". Chapter 25 budgets 2kb of script for two keyboard shortcuts, / for search and t for the file jump, and chapter 34 tells users to press them.

Both cannot hold. The policy now includes script-src 'self'. Inline script is still blocked, eval is still blocked, and nothing loads from another host, so the rule the policy existed to enforce is intact: keys.js is 814 bytes and a framework cannot arrive through it. TestScriptBudget fails the build if it passes 2kb or if a second script file appears.

29. Search has no index, on purpose, for now. Chapter 17 wants an index built on push and kept in the cache directory. This runs git grep over the repositories the asker may read, and no others.

Chapter 17's warning is that filtering a shared index after ranking leaks the existence and count of private matches, and calls that the highest-severity mistake available in the codebase. Not looking at all is the same rule applied one step earlier, so that class of bug cannot occur here.

The cost is that a query is O(readable repositories). An index has to keep this property when it arrives.

30. Copying with alternates makes a copy that a delete can destroy. Chapter 21.1 says to copy server-side using git alternates, which is right: it keeps a 4 GB repository off a home connection. It did not say what happens next.

A copy made with --shared borrows the original's objects. Delete the original and the copy loses the history it never had its own copy of, which turns "delete my repository" into "delete somebody else's work".

A delete now detaches its dependents first: git repack -a -d writes every borrowed object into the copy and the alternates file goes. A delete that cannot detach a dependent fails instead of proceeding. Chapter 21.1 says so now, and TestCopyAndDetach deletes the source and checks the copy still has its history.

31. A repository could be taken and never put back. Chapter 40.3 tells a user to move hosts with git push --mirror. Writing the test in chapter 45.1 showed that push being refused, ref by ref:

! [remote rejected] refs/meta/counter    (pre-receive hook declined)
! [remote rejected] refs/notes/runs      (pre-receive hook declined)
! [remote rejected] refs/proposals/2     (pre-receive hook declined)

The access matrix in chapter 18 says those namespaces are the server's or nobody's, which is right for a contributor and wrong for the owner restoring their own repository. Chapter 3's promise is that you can point your clone at another host and keep working, and half of it was missing.

The matrix now says the namespace owner may write any ref in their own repository. It grants nothing that was withheld: an owner who wanted to forge a build result could already push any content they liked.

32. Allocation could hand out a number already in use. Chapter 35.5 tells a user they can open a thread by pushing refs/notes/threads/<n> from their clone. That leaves refs/meta/counter behind, and the next proposal took the same number, putting two conversations in one place. Allocation steps over any number that already names a thread or a proposal.

Both were found by writing chapter 45.1's test, on its first two runs.

33. Anyone could take over anyone's proposal. Chapter 12 says only the proposal's author and accounts with push access may update its ref. The check read the commit author, git log --format=%an.

A commit's author is whatever the pusher typed into git config user.name. Setting it to lisa was enough to force-push over lisa's proposal. Writing chapter 45.2's hook tests surfaced it: the rejection said "belongs to tester", which is the name the test harness commits under, not the account that pushed.

The author is now read from the thread's meta blob, which records the account that pushed. Chapter 12 says so, and TestCommitAuthorIsNotIdentity performs the takeover and expects it to fail.

The two also differ in ordinary use: applying somebody's patch and pushing it is normal, and it should not hand them your proposal.

34. The chapter 25 budget is not met, and the reason is process spawn. Chapter 45.5 says to assert the budget against a large repository rather than a toy. Against 1000 files and 200 commits, on an ordinary laptop:

page before caching after caching budget
file tree 82ms 38ms 10ms
log with diffs 56ms 56ms 20ms
file view with blame 167ms 47ms 20ms

Payloads are all inside budget: 1kb, 18kb, 1kb against 15kb, 30kb, 40kb.

Caching blame by blob hash and the per-entry log by tree hash, both of which chapter 25 specifies, took the file view from 167ms to 47ms and the file tree from 82ms to 38ms. Neither reaches the number.

The floor is process spawn. git rev-parse HEAD on a small repository measures 8ms, almost all of it spawn. A page that runs four git commands has spent 32ms before rendering anything. Chapter 25 now states this and says the budget is a budget on git invocations as much as on milliseconds: one or two per page, reached with git cat-file --batch and one git log --patch for a whole page, and libgit2 in-process where that is not enough.

Cutting invocations closed most of it. git cat-file --batch answers the hash, the size and the content in one process instead of three. HEAD and refs are plain files, so reading them costs nothing where asking git costs 8ms each. Parsed configuration caches by the commit it came from, which chapter 14 asks for. The per-entry tree log keys on commit and path rather than tree hash, because the commit is a file read and resolving the tree is a process.

Each commit's diff is also cached by its own hash. A commit cannot change, so a push invalidates one entry rather than the page. The commit list comes from one cheap git log with no patch; on a miss the diffs come from a single git log --patch for the range rather than one call per commit.

page first measured now budget
file tree 82ms 10.3ms 10ms
log with diffs 56ms 24ms 20ms
file view with blame 167ms 10.5ms 20ms
threads list 104ms at 6 threads 55ms at 50 10ms
one thread not measured 22ms 10ms

Three pages read git once per row, and none of them were measured. The threads list read each thread's meta with a cat-file, listed each note tree with an ls-tree, and read every comment blob with another cat-file. Fifty threads with three replies each is roughly three hundred processes. The runs page did the same over the runs ref and then spent a git log per row for the commit subject. A single thread page did it over one thread's notes.

All three now use gitx.Batch, which drives one cat-file --batch process for many objects, and gitx.TreeEntries, which reads the raw tree git answers with. The threads list costs three processes whatever the thread count, the runs page three, and one thread two. Measured on the live server: threads list 104ms to 32ms at six threads, runs page 78ms to 32ms at four runs, one thread 32ms to 21ms.

The remaining cost is git working, not forge spawning. gitx.Run for a rev-parse measures 7.8ms on this machine and a bare exec.Command measures 9.1ms, so the wrapper adds nothing and the book's 8ms figure holds. Three processes is 24ms of that 55ms; the rest is git reading fifty trees and a hundred blobs, which is work no amount of batching removes. Chapter 25's 10ms is not reachable from a process, and the chapter already names the answer: libgit2 in-process rather than a looser number.

Getting the threads list to the chapter's stated two processes means dropping the for-each-ref and reading the ref files directly, the way HEAD is read. That is left alone deliberately: a direct read is wrong inside a worktree, which is a bug this codebase has already shipped once, and it saves 8ms against a page that misses by 45.

TestBudget is left failing rather than adjusted, because chapter 45.5 asks for build-failing thresholds and a threshold quietly raised to match the code measures nothing.

Also changed

SQLite or PostgreSQL. One setting picks it:

[database]
url = "sqlite:///var/lib/barerepo/forge.db"

SQLite is the default and is the right answer for almost every install: the server-owned data is small, writes are serialised by one process, and the backup becomes a file copy. Postgres exists for people who already run one. It does not make forge faster; the hot path is git.

Chapter 41.7.1 covers it. [paths] db is gone, replaced by [database] url.

Tests run on SQLite. It needs no service, so every test gets a fresh empty database. The Postgres schema is held to the SQLite one by a test that compares the two definitions column by column and needs no server. Run the full suite against a real Postgres before a release, not on every commit.

The write path was the last one starting a process per row

Reading a thread was batched; writing one was not. Write is a compare-and-swap on the notes ref, and the loser reads again rather than dropping a comment. That is correct, and under load it was quadratic: twenty writers meant the twentieth lost nineteen times, and each attempt cost a ls-tree plus a cat-file per note. Chapter 45.3's thousand comments took 323 seconds and came within one retry of the twenty the loop allows, which is where a comment starts being lost for real rather than in theory.

Two changes. Writers in one process now queue on a striped mutex, so the swap is contention between processes and not inside one. The read inside the loop uses gitx.Batch, so a thread costs two processes whatever the note count. The same test now takes 71 seconds, and nothing is dropped.

Serialising in one process leaves the swap untested by the tests named for it, because they can no longer collide. TestConcurrentRepliesAcrossProcesses runs eight real processes at one thread, which is the case that actually happens: a hook and the web server writing at once. Deleting the swap's old value makes it keep one comment of eight, so it is testing what it says it is.

countComments was dead and is gone.

The whole suite went from 232 seconds to 66, and three of the five chapter 25 budgets now pass that did not. TestBudget still fails on the threads list, one thread, and the log with diffs, and is still left failing rather than adjusted.

Three pages put the wrong thing at the right of the tab row

repotabs ended every page with "jump to file t". The mockups end four pages differently: threads with open · merged · closed · all, runs with runners · add a runner, runners with add a runner, releases with newest first. Only the code pages get the file jump, which is the only place the t shortcut goes anywhere.

The threads one was not a missing decoration. Chapter 24's thread list says "Filters are open, merged, closed, all" and forge had no filter at all, so a repository with three open threads and forty-one closed showed forty-four rows and no way to narrow them. The filter is a query parameter, the default is all, and an unknown value shows everything rather than an empty page. Abandoned counts as closed, because the row names four states and the model has five.

The counts in the header stay counts of every thread. A number that moves with the filter is the filter read back, and says nothing.

Still missing on the thread list row

Chapter 24 says each row shows "whether a ref is attached, diff stats if so, build status if any, and reply count". Forge shows the ref and the count. The mockup's has proposal · +81 -12 · build ok · 2 replies is two thirds there.

A diffstat per row is a git process per row, which is the thing chapter 25 forbids, so it wants the same treatment the commit diffs got: cache by the pair of end hashes, which are immutable, and pay only for proposals not seen before. The build status is already in the runs notes the runs page batches.

The log page now opens one diff at a time, against what the book says

Chapter 24's repository log says "the diff already expanded", chapter 34.1 says "Each commit shows its diff. Large diffs are collapsed. Press expand", and docs/BUILD.md repeats it. The page now ships every diff shut, and opening one shuts the last.

This is the author's call and it overrides the three places above, which should be amended. The reason it was asked for is visible in what the page actually did: twenty commits, thirteen of them over the inline threshold, so thirteen rows said "large diff collapsed" and seven showed a wall of diff. Which state a row was in depended on a size threshold the reader cannot see, there was no way to shut a diff once open, and "expand" was not an expand at all, it navigated to the commit page. Every complaint in that sentence is true whichever default is chosen.

No JavaScript was added. <details name="log"> is an exclusive group in HTML: opening one closes the others, with no script at all. Chapter 25 budgets 2kb of JavaScript for two keyboard shortcuts and this spends none of it. Where a browser does not know the name attribute it ignores it, and the diffs are still collapsible, just not exclusive, which is the right way for it to fail.

A large diff is still not inlined, because the payload budget is on what is sent, not on what is displayed, and a shut <details> has still sent its contents. Those rows say "too large to inline" and link to the commit, and they carry no disclosure triangle, because there is nothing behind them to disclose.

Things that should have been links and were not

The month on the keys and profile pages was formatted with "jan 2006". Go's reference month is Jan, so a lowercase one is not a token and was copied out literally: every key ever added read "added jan 2026" whatever month it was. Now formatted with the reference layout and lowercased afterwards, which is what the mockups show.

Then a sweep of every view for values that name something forge can open:

view was plain now opens
log the commit hash, the author the commit, the account
commit the parent hash, the author, each path, the branch the commit, the account, the file, the log
thread the merged commit, the proposal ref, each author, a line anchor the commit, a compare, the account, the file at the line
thread list the merged commit, the author the commit, the account
runs, run the ref, the runner, the subject, the run's hash a compare, the runner list, the commit
releases the tag, the tagger the files at the tag, the account
compare each path the file on the b side

A name in a commit or a pushed note is not an account. It is whatever the writer set in their local git config, so linking it blindly makes a page full of 404s. Every name is checked against the accounts table first, in one query per page rather than one per row, and a name that is not an account stays plain text. On the live server this is visible: lisa, dave and rock link, and donuts-are-good does not, because only the first three are accounts.

The same rule applies to hashes and paths. A deleted file gets no link, because the file does not exist at that commit; parsePatch now records Gone from git's "deleted file mode" line. A thread with nothing merged gets no merged link. An anchor whose original is lost gets no file link.

Every link on the log, thread list, thread, runs, releases, commit and file tree pages was then fetched. None of them 404s.

A readme link, and refs you can see before you type them

Two gaps the book does not name, both asked for directly.

The readme. Chapter 24 is against a rendered readme on the landing page, because that is the space GitHub spends instead of answering what changed. It is not against a repository saying it has one. The log bar now carries a readme link when the root tree holds one, under any of the usual names, and it opens the file view, which shows source with a blame gutter like every other file. Nothing is rendered on the log page and nothing moved off it.

Finding it costs one tree read, cached by the commit, so it is free after the first hit. The log page measures 16ms against chapter 25's 20ms with it in.

The refs. Chapter 24 says the compare fields are free text and gives the reason: master...refs/proposals/47 is a legitimate thing to type and a dropdown of branches cannot express it. That reasoning holds, and it left a repository's branches, tags and proposals invisible everywhere in the web interface. A field whose valid values cannot be discovered is a field only its author can use.

So the compare page lists what exists, under the form, as links that keep the a side and swap the b side. It is not a picker: the fields stay free text and the list is what is there to type. The branch name in the log bar links to that page, which is the only place a branch name meant anything before.

Notes refs are left out. They are storage, not somewhere to compare against.

The thread list row is now the whole row chapter 24 asks for

Chapter 24 says each row shows "whether a ref is attached, diff stats if so, build status if any, and reply count". Forge showed the ref and the count. Now it shows all four, in the mockup's order: has proposal · +2 -0 · build ok · 3 replies, with a failed build in the one colour the mockups use for it.

Neither number costs a process per row. Every run in the repository comes back in the two processes run.Recent already used for the runs page. The diffstat is keyed on the two end hashes, which are objects and cannot change, so a proposal is measured once and never again. Resolving a proposal ref to a hash is a file read.

The threads list measures 41ms against chapter 25's 10ms, up from 35ms, on a cold cache in a repository built by the test. It was already failing that threshold for the reason SPEC-NOTES gives above: git reading fifty trees is work no batching removes.

Two bugs from one wrong assumption about HEAD

gitx.ResolveRef read a ref file and returned its contents. For every ref under refs/ that is a hash. For HEAD it is the line ref: refs/heads/master, because HEAD is symbolic.

So ResolveRef(dir, "HEAD") returned a ref name where every caller expected a hash, and returned it with no error, which is the worst shape a wrong answer can have. It broke two things in one afternoon: the readme page 404'd, because the name is not a tree, and every proposal's diffstat came back +0 -0, because the name is not a rev.

Fixed in gitx rather than at either call site, since the next caller would have hit it too. ResolveRef now follows ref: up to five times and errors on a ref that points at itself, which is a file somebody wrote by hand.

Seventeen templates were never rendered by the test that claims to render them

TestTemplatesRender opens "Every template runs here, because a typo otherwise fails when a user asks for the page." It ran ten of twenty-seven. The other seventeen, including every account page, the runner setup, search and the inbox, had nothing rendering them at all.

The test now walks pages and fails on any template with no case, so the next one cannot ship untested. The seventeen have cases.

Also

The log rows are uniform: every commit carries a view commit link, and the "too large to inline" wording is gone. A link inside <summary> navigates without toggling the disclosure, which was checked in a browser rather than assumed.

The rendered readme keeps each control once: rendered in the bar, [source] under the prose, raw in the footer.

The breadcrumb named two places and linked neither

rock / forge / 26ddf1d sat at the top of seventeen pages with the account and the repository as plain text. Both are places, and the reader is one click from each on every page except the one that made them type the URL.

There is now a crumbs template, so a page cannot write the pair by hand and forget, and a test fails on any template that goes back to doing so. The last segment stays plain, because it names the page you are already on.

The file tree is the one page where the middle segment moves: at the root the repository is the page, and under a path it is somewhere to go back to, so it links only in the second case.

The one line on the runner page pointed at a 404

BUILD.md calls the runner command "the piece to get exactly right, because it is the most visible proof of the no-interstitial thesis". The page rendered it correctly and both halves of it, /runner.sh and /runner.ps1, returned 404. So the page that exists to prove one paste is enough handed out a paste that failed.

The protocol and the forge runner subcommand were already built. Only the script was missing.

Where the binary comes from. BUILD.md says "same single binary as the server, different subcommand", so the server hands out its own executable at /runner/binary, guarded by the runner token, which the script already has and a stranger does not. No release hosting, no second artifact to keep in step.

What it does when it cannot help. A server can only serve the platform it was built for. The script compares uname against the server's and, on a mismatch, prints the two commands that build and run it instead. A wrong architecture installed silently is worse than a refusal that says what to do.

Measured on the live server, the real line: downloads 27mb, attaches, and the runners page shows the machine as idle. That is chapter 15's "the page the user copied from updates the moment the runner attaches", working.

The book writes the url as the hosted barerepo.sh. Serving it from each install's own external_url is what a self-hosted forge can actually do, and it is what the template already rendered.

Reading GitHub workflows, and the book grew a chapter for it

The book had nothing on this. It has chapter 15A now, and BUILD.md has a stage 6A, because the author asked for the feature and a feature the book does not describe is a feature nobody can check.

It is not a contradiction. Chapter 3 already credits .github/workflows with learning the lesson forges forgot: configuration is a file in the tree. The objection in this book was never to that file. It was to a forge configurable only through its own web forms. Reading a workflow honours the same rule .barerepo/config honours.

What it translates. Every run: step into one shell script, in order, with env: turned into exports and working-directory into a subshell so a step's directory does not leak into the next. actions/checkout is answered rather than run, because forge cloned the repository already. container: becomes the image.

What it declines, and the rule the feature rests on. Any other action, any if:, any matrix, any shell forge cannot start, and any ${{ }} left in a command. Each one is printed in the push output, naming the step and the reason. A skipped step is never silent. A build reporting success while quietly running half of what was asked is worse than no build, and it is the failure every partial implementation of somebody else's format drifts toward.

runs-on, against machines that actually exist. A job now carries the labels it asked for, and a runner takes the first queued job it satisfies rather than the first queued job, so a build waiting for a machine nobody attached does not block the ones that could run now. ubuntu-latest and its siblings match on the operating system the runner reported when it attached, self-hosted is always true because every forge runner is, and anything else must be a label the runner declared.

When nothing fits, the push declines and prints the page that fixes it, which is chapter 11's rule about typos applied to machines:

unit wants a ubuntu-latest machine and none is attached.
  attach one: https://barerepo.example/john/johnbot/runners/new

Precedence. [build] command wins and the workflow is not read at all. A repository that answered in forge's own file is not second-guessed, and a repository with both does not build twice. Both directions are tested.

The one dependency. gopkg.in/yaml.v3. Hand-rolling a YAML subset for a format other people control is how a parser becomes a bug farm, and the book's minimalism is about architecture, not about refusing a parser.

The jobs table grew a labels column in both dialects, which the portability test compares column by column and passed.

Making a workflow work without editing it, and a silent pass it turned up

The aim the author set is that a repository arrives with the file it already has and builds. Two things stood between that and the first implementation.

${{ }} was not a decline, it was a break. The first version left an expression in the command and reported it. sh reads ${{ as a bad substitution and fails the step outright, so any workflow naming ${{ github.sha }} failed at the first step that used one. That is the most common thing in a real workflow after run: itself.

Forge now sets the environment GitHub sets, including GITHUB_SHA, GITHUB_REF_NAME, GITHUB_REPOSITORY, GITHUB_WORKFLOW, GITHUB_JOB, CI, and the runner's own RUNNER_OS, RUNNER_ARCH and RUNNER_TEMP, and fills the expressions that name the same things. A workflow reading either form now works untouched. This is a substitution and not an evaluator: secrets, and anything else outside the list, is still left alone and still reported.

A failing step did not fail the build. The runner passes the command to sh -c with no set -e. For chapter 15's one command that is correct, and its exit code is the command's. For a translated workflow of many steps it is not: the first step could fail, every later step still ran, and the build reported the exit code of the last one. Measured: a script whose first line is false exits 0.

That is precisely the green build that means nothing, which is the failure the whole chapter is written against, and the feature shipped with it. The generated script now begins with set -e, which is also what GitHub does. [build] command is untouched, being one command by definition.

The setup actions check, they do not install. setup-go, setup-node, setup-python, setup-java and setup-dotnet become a test that the tool is on the machine, failing with a sentence if it is not, and printing the version the workflow asked for beside the version the machine has. The version is printed rather than enforced, since a patch digit should not fail a build.

Not installing is the point. The runner is a machine the user owns, and a build that quietly puts a toolchain on somebody's laptop is doing something the person who pasted one line did not agree to. Chapter 15 says a build has access to whatever the machine has. It does not say a build may change what the machine has. A test asserts no generated line runs apt-get, brew install, npm install, pip install or a piped shell script.

actions/cache is answered as a statement that forge does not cache, so a build starting from the clone is a stated fact rather than a surprise.

A matrix builds a job several times, so forge queues it several times

The first version declined a matrix whole, which meant a repository whose CI is a matrix got no build at all. That is most repositories that test more than one version of anything.

The axes are multiplied, exclude removes what it names, and each combination becomes its own job with its own runs-on. The combination is substituted into the command, the image and the machine, and exported as MATRIX_<AXIS> so a step reading the environment works as well as a step reading the expression.

include is not applied. It can add keys to a combination and whole combinations no axis names, and a wrong guess there runs a build the workflow did not ask for. Forge says so and builds the axes.

The part that made it usable rather than merely present. A matrix queues one commit several times, and run.Record had no way to say which build was which: three combinations produced three records distinguishable only by the machine, and not at all when two ran on one machine. A matrix build where you cannot see which combination failed is not worth having.

The job now carries a name, test (go 1.26, os ubuntu-latest), with the axes named and sorted so the same matrix always produces the same names. The server already holds the job when the run finishes, so the name reaches the record without the runner protocol changing at all. The runs page and the run page both show it.

The jobs table gained name beside labels, in both dialects, which the portability test compares column by column and passed.

What this looks like when the machines are not all there. A matrix over ubuntu and macos, on a server with only a linux runner attached, queues the ubuntu half and declines the mac half by name, with the link to attach one. That is tested end to end, along with each machine taking only its own combination.

The runs page said builds were off while builds were running

BuildOn was [build] command != "", which was the whole truth until chapter 15A gave a repository a second way to build. A repository building from .github/workflows showed "builds are off. set [build] command to turn them on" underneath its own build results.

The page now says what the repository actually builds from, naming the files, and the off state names both routes rather than only forge's own. Finding out costs one process and only when [build] command is absent, since a repository that answered in forge's file is not asked a second question.

This is the same shape as the seeded runs that read "builds are off" earlier in this file: a page holding two facts that contradict each other, where one of them was written before the other existed.

I added two columns without a migration, and every existing install would have broken

The jobs table gained labels and then name for chapter 15A. I added them to the CREATE TABLE in migration 5, which is the one migration this codebase already says must never be edited: migrations is append-only and schema_migrations records how far an install has got.

A fresh database therefore had the columns and every existing one did not. The first /runner/poll after the upgrade answered no such column: labels, 500, forever. Builds would have stopped on every server that had ever run forge before, and only on those.

Nothing in the suite could have caught it. BUILD.md says tests run on SQLite because it needs no service, so every test gets a fresh empty database. That is the right call and it makes the whole suite structurally blind to an upgrade. It was found by pasting the runner line at the development server, whose database was three days old.

Fixed by putting migration 5 back as it was and adding migration 6, which is what the mechanism was always for.

Two tests now cover the class:

TestAnExistingDatabaseUpgradesToEveryLaterVersion opens a database as an older forge that knows only migrations 1..n, closes it, reopens with the full list the way a replaced binary does, and then runs the real query. It does this for every n, so a column added without a migration fails at whichever version predates it.

TestMigrationsAreAppendOnly holds the count. Adding a migration means raising one number and is meant to be easy; editing an earlier one changes the count in the other direction and fails with the reason spelled out. Falsified by putting the original mistake back, which it catches.

The runner needed a flag the book says it does not

Appendix E gives forge runner <token> [--labels a,b], chapter 38.2 gives forge runner rt_live_7Kq2mXe --labels build,test, and the runner-setup mockup gives the same. Forge's own page printed --server https://... in the middle of it, and the binary refused to start without it.

A token that cannot say where it came from forces a second parameter, which is the interstitial chapter 15 exists to remove.

The runner now records where a token attached, after the attach succeeded so a wrong address is never the one kept, in ~/.config/forge/servers. The file maps a hash of the token to a url and holds no token, at mode 0600. The pasted line supplies the address the first time and the documented line needs none after.

Measured on the development server: the one-liner attaches, and then forge runner <token> --labels build,test attaches with no address at all.

A matrix could show a green thread row over a failed build

The thread list reads the runs for a proposal and stopped at the first one that named the right ref. That was correct while a proposal had one build. A matrix gives it several, run.Recent returns them newest first, and the newest is not the one that matters.

So a proposal whose linux build failed at 10:00 and whose mac build passed at 10:01 said build ok. Chapter 19.1 says green is not news; a green that is hiding a red is worse than not news.

The row now reduces every run for the ref to one status, and any failure is the status: 1 of 3 builds failed, in the colour the mockups keep for it. One run still reads exactly as the mockup does, build ok · test ok, because that is what the build reported and there is nothing to summarise.

The thread page did not have the bug, since chapter 24 puts each build in the timeline as its own event and it never stopped early. It did drop the job name, so three builds of one proposal read as three identical rows. They now say which combination each was.

The escape hatch was written and never wired

internal/webhook held the deny list, the HMAC signature and the retry, and nothing in the tree called any of it. A repository could put [[webhook]] in .barerepo/config and forge would do nothing with it, without saying so.

Chapter 23.1 calls a webhook the escape hatch, the thing that makes it acceptable to refuse every integration request forever. An escape hatch that is present in the source and absent at run time is worse than one that was never started, because the config file accepts the lines.

Where it fires from. The event log, after a cursor, read by the server process. Both writers of events keep working the way they did: the hook process records a push and returns, the web process records a comment and answers. A receiver that takes thirty seconds cannot slow a push, because nothing on the push path is waiting for it.

A first start puts the cursor at the newest event. A new hook is not a request for ninety days of history.

The secret is a name. secret_env = "DEPLOY_HOOK_SECRET" names a value in the server's environment. A name whose value is not set is a recorded failure that says which name, not an unsigned delivery to a receiver expecting a signature. Committing a secret to a public repository stays impossible by construction, which is the whole reason 23.2 names it rather than holding it.

Twenty failures in a row stops a hook, per 23.4, and the config page is the only place that could report it, so it does: the url, the events it asked for, and either when it last delivered or the reason it is failing. One delivery clears the count, or a hook that fails once a week eventually stops for nothing.

Proven on the development server rather than only in tests: a pushed config naming a hook with an unset secret produced

1 failure since the last delivery · DEPLOY_HOOK_SECRET is not set on this
server, so nothing would sign the body

One dead receiver held up every other repository

The first sender walked one event at a time and each hook in turn, three attempts with backoff on every one. A receiver that is down costs about six seconds per event and twenty events before it disables, and nothing else in the queue moves for those two minutes.

A hook that has already failed now gets one attempt. The three attempts are for a receiver that is briefly down, and the first failure is the test of that. One event's hooks go at once, since they are independent of each other.

The budget was failing and the fork was most of it

Chapter 25's thresholds are asserted, and two rows had been red: the threads list at 39ms against 10, and one thread at 12ms.

Measured rather than guessed. Every git process this machine starts costs 6.5 to 8ms before git does any work, which the file tree page shows exactly: one ls-tree, 6.4ms, and 6.7ms on the clock. The threads list started four processes.

git cat-file --batch is now kept open per repository and pooled, so reading objects starts nothing. One thread went from 12ms to 400 microseconds.

I guarded against something that does not happen. A long-lived reader cannot see an object that arrives in a new pack, I assumed, so I stat'd every pack directory and restarted the process when it changed. It was wrong: git rereads the pack directory every time it answers missing, so the kept reader does see it. The guard is deleted. The test that made me delete it stays, because that reread is the assumption the entire pool rests on, and a future git that stops doing it must fail here rather than serve a page with a hole in it.

And a deadlock that does happen. The batch wrote every object id before reading any answer. Past the pipe buffer that is a deadlock: git stops reading input while it is blocked writing output, and forge stops writing while it is blocked writing input. Four thousand ids reproduces it. The write moved to its own goroutine. Both tests were falsified before being kept.

The threads list itself. It read every comment blob of every thread to count replies, and forked for-each-ref to find them. Refs are files, which is chapter 6, so the ref list costs no process now either. A row is cached by the note commit that wrote it, and a commit is immutable, so the entry never needs invalidating. The page reads nothing it has read before.

39ms to 7.4ms. The whole budget passes.

A url the book hands out, that answered 404

Chapter 19.5 lists four feeds, chapter 39.4 tells a reader to paste two of them into a feed reader, and appendix C lists all four as routes. Three were served. /<user>/<repo>/threads.atom fell through the route table to the 404 page.

The book prints that url twice, so the failure is not a missing feature, it is a promise the running server does not keep. A reader who follows 39.4 gets a feed reader with a dead entry in it and no reason given.

It now answers, with the same events the threads page shows and nothing else: the six thread and proposal kinds from 19.1. A push is not discussion, so the repository feed carries it and the threads feed does not. Both are asserted against a live server rather than against the query, because the route was the part that was missing.

A private repository answers 404 on both feeds, since 19.5 says public only and existence leaks.

The profile page had the same shape of gap. profile.html puts atom in the sidebar under the key count, and /<user>.atom has worked since the feeds went in, but the template never linked it. The link is back. A feed nobody can find from the page is a feed that needs the book open next to it.

The older link was drawn and never given a value

repo-log.html puts older in the footer, and the template has carried {{if .Older}}<a href="{{.Older}}">older</a>{{end}} since the log was built. Nothing ever set Older. The condition was false on every page forge has served.

So the log showed twenty commits and the twenty-first was unreachable. A repository with two hundred commits published a hundred and eighty of them over git and none of them over http.

The page now starts where ?from=<sha> says, asks for twenty-one, and links the twenty-first as older. There is no offset and no cursor to keep, because a commit already names its own position in the history. A from that is not a commit here answers 404 rather than quietly showing the newest page, since a stale link that looks like it worked is worse than one that says it did not.

A tag cost a process, and a hundred tags cost a hundred

The releases page called git notes show once per tag, inside the loop over for-each-ref. Measured against chapter 25's ten milliseconds, with a hundred tags: 1.785 seconds, and 29kb over a 15kb payload budget.

Three things were wrong and each is worth naming.

The notes. They are now read the way chapter 16's other note refs are: walk the notes tree through the object pool, one level of fanout at a time, skipping any subtree that holds no note this page asked for, then read the bodies in one batch. No process at all, and only the twenty bodies the page draws.

for-each-ref. Eleven milliseconds on its own for a hundred tags, which is the whole budget. Refs are files, per chapter 6, so the tag names and their object ids come from the ref files, and the tag objects come from the pool.

Batching by ref name still cost 11ms because git resolves each name; batching by the object id the ref file already gave costs 7.7ms. Then the sorted list is cached under a digest of the ref state, so a page that follows a page with no tag pushed between them reads no objects at all. The key is the content, so nothing invalidates it.

No end to the list. Twenty rows, then older, the same word in the same corner as the log. The releases mockup has two releases and so shows no such link; the log mockup does, and one convention for "this list continues" beats inventing a second.

1.785s to 7.4ms, 29kb to 6kb. That row of the budget passes.

Still failing, and it was failing before this pass. The log with diffs takes 26ms against 20. Measured at the previous commit, with none of this pass's changes and no tags in the repository, it took 24ms. The file tree sits on 10ms against 10 and crosses in either direction between runs. Neither is caused by anything here, and neither is fixed by anything here.

The landing page spent nine milliseconds asking whether the repository was empty

Chapter 25 gives the log with diffs twenty milliseconds. It took twenty-six, and the last pass could not say why. This pass measured the parts instead of reading the code, which is the method that worked on the releases page.

page 22.9ms
  gitread.Log     12.3ms, of which one fork is 11.5ms
  transport.Open   0.2ms
  readme, config   ~0ms
  git --version    8.2ms

That last line is the important one. A bare git --version costs 8.2ms on this machine, so a process is 8ms before git does anything. Two processes is 16ms of a 20ms budget.

The log page started two. The second was repo.IsEmpty, which ran git for-each-ref --count=1 on every repository landing page to answer "does this repository have a ref". Refs are files, per chapter 6. gitx.AnyRef walks refs/ and stops at the first one, then falls back to packed-refs, which is where git gc moves them. No process.

26ms to 13ms. What remains is the one git log, and that one has to be a process, because the order of a log is a revision walk and forge is not going to reimplement one.

The file tree ran ls-tree for a listing the object pool already had

The same measurement put the file tree on 10.2ms against a 10ms budget, which means it failed about half the runs. One ls-tree --long fork was the whole of it.

--long is there for blob sizes. Nothing displays them: not the template, not the mockup, and Entry.Size had one writer and no reader. So the size was the only reason to ask git rather than read the tree object, and the size was never used.

gitx.TreeRows reads a raw tree object, keeping the mode, since the mode is the only field that says tree or blob. Entry order is git's own tree order, which is what ls-tree was printing anyway, so the directories-on-top pass below it still sees exactly what it saw.

10.2ms to 0.5ms.

Rendered diffs are cached now, which the build guide asked for

BUILD.md says to cache rendered diffs because they are immutable. Forge cached the patch text and re-parsed and re-rendered it on every request.

The hunk markup was written three times, identically, in repo-log.html, repo-commit.html and repo-compare.html. It is one hunks block in the layout now, and the log renders its commits through it once and keeps the html under the commit hash.

Worth saying plainly: this was worth 0.7ms of the 10ms I thought it would fix. I had assumed the render was the cost and it was not. The measurement above is what found the two forks. The cache stays because the guide asks for it by name and because it took triplicated markup down to one copy, not because it was the fix.

The whole budget passes, with margin on every row:

file tree             0.5ms      1kb   (budget 10ms, 15kb)
log with diffs       13.1ms     22kb   (budget 20ms, 30kb)
file view with blame 10.3ms      1kb   (budget 20ms, 40kb)
threads list          2.1ms     13kb   (budget 10ms, 15kb)
one thread            0.6ms      1kb   (budget 10ms, 15kb)
releases              7.5ms      6kb   (budget 10ms, 15kb)

Copying was written, and had no way to reach it from the site

repo.Copy clones with --shared, fetches notes and the counter, drops proposal refs and installs hooks. repo.Detach un-borrows. repo.Dependents finds who borrows. The delete path already calls Detach before it trashes anything, which is the requirement in chapter 21.1 that stops a copy losing its history.

forge copy <src> <dst> uses all of it. Appendix C lists POST /<user>/<repo>/copy and nothing answered it. The escape hatch shape again: the machinery was complete and one route was missing.

Who gets the button. Not the owner. Chapter 21.1 exists because chapter 12 removed the fork, and the thing being restored is taking a project somewhere its maintainer will not go. So the control is on the config page for any signed in reader who can read the repository. Read access is the whole permission, because chapter 12 already removed asking as a step.

Where it goes. <you>/<the same name>, with no field to fill in. The book's own example is forge copy john/johnbot lisa/johnbot. If you already have a repository by that name the page says so and links it, rather than offering a button that will fail.

What it says. The same four sentences the CLI prints, because the terminal and the page must not explain the same operation differently: branches, tags, history, threads and notes come across, proposal refs do not, there is no link back and no badge, and contributing means pushing a proposal.

The plain commands come first, per chapter 5 rule 2. git clone --mirror, then a push naming heads, tags and notes, which is exactly the ref set the server side copy moves. Forge creates the destination on push, per chapter 11, so the plain path needs no visit to /new.

Proven end to end against a running server: lisa copies john's repository, the copy has master and does not have the proposal ref that was pushed to the original first, and john deleting the original leaves lisa's history intact. The test asserts the original had a proposal ref before the copy, so the assertion that none came across cannot pass by accident.

The last url in appendix C that answered 404

GET /<user>/<repo>/release/<tag> is in the route table. Chapter 24 has one entry for releases and it describes the list. There is no mockup for a single release and no paragraph describing one.

Building a page the plans do not describe would be inventing product voice, which is the mistake that produced a landing page full of made up copy earlier in this work. Answering 404 to a url the book prints is the mistake fixed two passes ago.

So it opens the list at that tag, using the ?from= the releases page already takes. A release is a tag, a body and some files, and all three are on that row. A tag that is not in the repository is a 404, not the newest release wearing the wrong name.

Every url in appendix C now answers.

A repository search result printed its own name twice

search.html gives one thing a box: the matched source line. A code row is the kind, the file and line, then the line itself in a box with the match marked.

A thread row in the mockup is johnbot 44 · does this work behind a socks proxy? on one line, with the excerpt underneath in muted text. A repository row is john / johnbot and its description underneath. Neither has a box.

Forge gave every row a box, filled with Result.Text. For a thread that put the title in a monospace code box. For a repository it put rock / forge in a box directly under the link that already said rock / forge.

The box is now the code row's alone. A thread carries its title beside its number where the mockup puts it, and a repository carries only its description. The query is marked in the heading line as well, since that is now where a thread title lives.

TestOnlyACodeSearchResultDrawsABox renders one of each kind and counts the boxes. A want list of strings cannot say "and not this".

The rule about comments was not being checked where comments also live

Every comment in this tree is one line. I had been checking .go and .css with an awk one liner and had never looked at .js, and the one script in the product opened with a two line block.

TestEveryCommentIsOneLine walks the tree and fails on any run of consecutive // lines in a .go, .js or .css file, and on a /* */ that does not close on the line it opened. It was falsified before it was kept: a temporary file with a two line comment fails it by name and line.

Search has no index, and chapter 17 opens by asking for one

Recorded rather than fixed, because it is the largest thing left and half of it would be worse than none.

Chapter 17: "One index, one result set". "Index on push, incrementally, from the pushed range rather than a full rescan. Index thread comments on note write." Appendix E lists forge doctor --reindex to rebuild it.

There is no index. readable opens every repository on the server for every query, and search.Search then runs git grep in each one, plus a thread read. Measured on the development server with 29 repositories: 64ms, against chapter 25's ten milliseconds for a page with no diff. It is O(repositories) in processes, so it gets worse with exactly the growth a forge wants.

What is already right, and must stay right when the index arrives: the read filter is applied before the search, not after. readable builds the target set from what the asker may open, so a private match is never ranked and then dropped. Chapter 17 calls filtering after ranking a leak of the count of private matches, and BUILD.md calls it the highest severity mistake available here.

The constraint that decides the design: BUILD.md requires a schema both SQLite and PostgreSQL accept, so FTS5 and tsvector are both out. One document table with a repository column, filtered in the query, is portable and is one query instead of N processes.

An empty package

internal/web was an empty directory imported by nothing. Deleted.

Search has an index now, and the first thing it indexed was the trash

Chapter 17 opens with "one index, one result set". Forge had no index. Every query opened every repository on the server and ran git grep in each one, plus a thread read. Measured on the development server with 29 repositories: 64ms, against chapter 25's ten for a page with no diff, and O(repositories) in processes, so it got worse with exactly the growth a forge wants.

The read filter is the whole design. Chapter 17 says filter in the query, and BUILD.md calls filtering after ranking the highest severity mistake available here. So every document carries public and a readers column holding the owner and [access] push as |john|lisa|, and the query is

WHERE (LOWER(body) LIKE ? OR LOWER(title) LIKE ?)
  AND (public = 1 OR readers LIKE ?)

The pipes matter: |john| never matches inside |johnson|. Chapter 18 makes read binary, public or the owner or [access] push, so the whole rule fits in two columns and needs no join.

That test was falsified before it was kept. Changing public = 1 to 1 = 1 makes it fail by name, for both an anonymous reader and a signed in stranger.

Why not FTS5 or tsvector. BUILD.md requires a schema both SQLite and PostgreSQL accept, so neither becomes the only one actually tested. Neither full text extension is portable. One document table with a LIKE scan is, and a scan of one table beats N processes by a wide margin, which is the whole problem being solved.

What is indexed. One row per file at the tip of the default branch, one per thread with its comments, one per repository for its name and description. Blobs over 512kb and anything holding a NUL byte are skipped: chapter 17 indexes source, and one generated file should not become the index.

Incremental, as the chapter asks. A push takes git diff --name-only old..new and updates only those paths, deleting the rows for paths the push removed. A first push, or a push whose old side is unknown, walks the tree once. Code lives at the tip of the default branch, so no other ref changes what a code search finds, but every push rewrites the read set and the discussion, because visibility arrives in the tree and a proposal push carries notes.

A comment written on the web is a note write and not a push, so the three places httpd writes a note reindex the discussion.

forge doctor --reindex is appendix E's line, and it empties the index first. A rebuild that only adds cannot remove a repository that has gone.

Which is how the bug was found. The first reindex on the development server said "indexed 7 repositories" and one of them was trash/1787104067-mark-renamed. Chapter 44.4 keeps a deleted repository for 30 days, and the walk was matching every *.git directory under the repository root, including the ones waiting to be erased. A deleted private repository would have had its contents searchable under an account named trash.

The delete path already dropped its documents, so this was reindex alone. The walk skips the trash directory now and the test that proves it was falsified first: without the skip it fails with the actual leaked link, trash / 1787130049-john-johnbot / retry.go:3.

64ms to 7.8ms on the development server. In the budget, over a thousand indexed files, fifty threads and a hundred tags:

search   900µs   1kb   (budget 10ms, 15kb)

Every row of chapter 25 still passes.

Giving a repository away left the old owner able to search it

A transfer moves the directory, the database row and the index rows. It did not move the read set. Chapter 18 makes read binary and the owner is half of it, so the readers column still said |john| after john gave the repository to lisa.

Two things followed, and both are wrong in opposite directions. Lisa could not search the private repository she now owned, which is exactly the failure chapter 17 names when it refuses to solve access by indexing only public work. And john could still find its contents, in a repository he no longer owned.

A rename does not have the problem, because a rename does not change the owner. The read set is rewritten from the tree after any move, since a transfer is the one change to who may read that no push announces.

The test fails on both halves without the fix.

A query is a string, and LIKE thinks some of it is syntax

% and _ are LIKE's wildcards. A search for 100% was reaching the database as %100%%, and a search for read_all matched readXall.

They are escaped now, with ESCAPE '\', which both dialects read the same way.

The first version of this test did not actually test it. %100%% still needs the literal 100, so it matched the one document it should have matched and the test passed with the escaping removed. It asks for c%e now, a string that appears in neither document and which unescaped reaches both through the word "coverage", and for read_all against a document holding readXall. Both halves fail without the escaping.

Worth writing down as a rule rather than an incident: a test that passes when the thing it tests is deleted is not a test. Take the code out and watch it go red.

One row per file hid the rest of the matches

git grep returned every matching line, so a symbol used four times in a file was four rows. The index holds one row per file, and the first version of the query turned that into one row per file on the page as well.

That is a loss for the one thing chapter 17 says a forge is usually searched for. A file now contributes up to three lines, each with its own number and link, and the whole set is still bounded by the page's limit, so a common word cannot fill the page from one file.

The tab strip carried a count that only one page filled in

Every mockup draws the tab row as log · files · threads 3 · runs · config. The number is the point of putting it there: a reader learns there is discussion without spending a page load to find out.

openRepoFor builds that row for every repository page and passed 0. The threads page overwrote it afterwards with the real number, so the count appeared on exactly the one page where a reader already had the list in front of them.

It is read once in openRepoFor now. thread.OpenCount goes through the same cached row list the threads page uses, so with fifty threads the cost is the file tree at 0.5ms to 2ms and the log at 12.9 to 14.5, both well inside chapter 25.

The thread page's footer had the same gap. thread.html in the mockups says threads · 3 open and the template said threads.

A profile row was missing the number chapter 24 asks it for

Chapter 24, profile: "each with language, size, default branch, and open proposal count". repoMeta built the first three. The mockup line is go · 4.1mb · master · 3 proposals and forge drew go · 4.5mb · master.

An open proposal is an open thread that carries a ref, which is what the threads list already means by the word, so the count comes from the same place. A repository with none says nothing rather than "0 proposals".

The bio in profile.html has nowhere to live, and that is correct

The mockup puts writes bots. mostly go. under the account name. There is no such field and there should not be one.

Chapter 10's closed list is the complete set of what the server stores outside git. It has six items and none of them is a profile. The chapter says outright that the list was four items in an earlier draft and grew without being updated, and that "a closed list that quietly grows is worse than an open one".

A repository description lives in .barerepo/config, in the tree, which is why that one is drawn. An account has no tree to put a bio in. Left unbuilt on purpose.

A note to myself about git checkout

Falsifying the tab count test meant editing web.go, running the test, and putting the line back. I put it back with git checkout internal/httpd/web.go, which threw away every other uncommitted change in that file from the same pass: the struct field, the profile count and the repoMeta signature. The test went red for the wrong reason and I had to write them again from the transcript.

Falsify by reversing the exact edit, never by checking the file out.

The keys page named one credential kind out of four

Chapter 24 is explicit about why the closing line exists: "The page names what each one can do, because a user about to paste a token into a feed reader deserves to know it cannot write."

Forge said "a key signs you in and pushes. a git token clones and pushes over https." and stopped. The two sentences it dropped are the two the chapter gives a reason for. keys.html has them both: "a runner token attaches one machine to one repository. a feed token reads one feed and can write nothing."

All four are named now. The runner sentence also says where a runner token comes from, because forge has no new runner token button and the mockup does.

That missing button is deliberate. Chapter 15 puts the token inside the copied command, per repository. A button on /keys has no repository to scope to, so it would have to ask for one, which is the interstitial the whole chapter exists to remove. The runners page of a repository is where the token is made.

A runner token said nothing about the runner

Chapter 24 asks a runner token row for "its labels, attached machine, and last-seen time". Forge showed the scope and the last-used time.

The runners table already carries token_id, so the machine that attached with a token is one join away. A row now reads runner · rock/forge · m2 · labels build, test · last used 2h, which is the mockup's john/johnbot · labels build, test · last seen 40m with the machine named as well.

Runners were never busy and had never run anything

Chapter 24, runners: "Attached machines, platform, labels, run count, status, last seen". The run count was absent and the status had two values where the mockup has three.

Both come from the jobs table, which already carries runner_id, in one grouped query: how many jobs each machine has taken, and whether any of them is running now. A machine holding a job reads busy, which is what runners.html shows for lisa-mbp, and a machine that has run nothing says nothing rather than "0 runs".

Pages checked this pass and found correct

The sweep is worth recording in both directions, or the next pass repeats it.

  • inbox matches inbox.html completely, including the horizontal rule at last_visited from chapter 19.4 and the ninety day line.
  • commit correctly has no tab strip, because repo-commit.html has none.
  • compare matches, including the ref shortcuts under the form.
  • file tree, runs, thread, releases, threads match.

Adding a runner had grown the second step chapter 24 forbids

Chapter 24 on the add-a-runner page: "Three commands, one per platform, all visible at once, each with the token already inside it." Then, in bold: "If this page ever grows a second step, something has gone wrong."

The page showed one button, new runner token. You clicked it and then got the commands. That is a second step, on the page the chapter picks out as the clearest demonstration of the whole thesis.

The token is minted on arrival now. The reason the button existed is real: a token is shown once, so a page that mints on every view leaves a dead token per view. So a view first revokes this account's runner tokens for this repository that are over an hour old and that no machine ever attached with. A reload cannot pile them up, and cannot revoke the line the reader copied a moment ago either, which is the case the test names.

What I did not build: the auto-refresh. The mockup's third fact is "this page refreshes the moment a runner attaches", and chapter 24 lists it. It cannot be done here without breaking something else. A meta refresh on a page that mints a token either mints one per tick or revokes the line the reader is in the middle of pasting. Forge's page does not claim to refresh, so nothing on it is untrue, and the honest fix needs the page to know a runner attached without reloading, which is polling, which is javascript this design does not want. Left out on purpose.

A new repository could not be given a description

Chapter 24, new repository: "Shows. Name, description, default branch, visibility, create." The form had name, default branch and visibility.

The reason it was missing is real: chapter 11 says the server commits nothing, so a description has nowhere to go. It lives in .barerepo/config, in a tree that does not exist yet. Visibility has the same problem and was already solved, by carrying the answer to the empty repository page and putting it in the block the reader pastes.

The description now goes the same way.

The paste block changes shape when there is one. The mockup writes the config with printf '[repo]\nvisibility = "public"\n', and printf reads backslashes, and the whole argument is inside single quotes. A description holding a quote or a backslash would break the command the reader pastes, silently, on their machine.

So a described repository gets a quoted heredoc instead, which passes every character through untouched:

mkdir -p .forge && cat > .barerepo/config <<'EOF'
[repo]
visibility = "public"
description = "irc bot that refuses to leave"
EOF

A repository with no description keeps the mockup's printf line exactly. The only character a quoted heredoc cannot carry is a newline, and a form input cannot hold one, but it is stripped anyway rather than trusted.

The push rejected page was never built, and it is the one that matters most

push-rejected.html is a mockup, chapter 24 has an entry for it, and BUILD.md stage 4 says "this page matters more than it looks, because rejection is where new contributors get stuck". There was no template, no route, and no url.

The wording had been done. An earlier pass made the hook print the mockup's exact sentences so the terminal and the page could not differ. The page they were copied from did not exist.

Chapter 24 says why it has to: "The hook already printed the reason. The page exists because terminals scroll, and because a rejected push is where a new contributor decides whether to keep going." And: "The hook prints a URL alongside the rejection message. Without that line the page is unreachable, because a rejected push happens in a terminal and no browser is involved."

It holds no state. A rejection is not on chapter 10's closed list and must not be, so nothing is written down. The url carries the one fact the server cannot recompute, which is the ref you pushed to:

http://barerepo.example/john/johnbot/rejected?ref=refs/heads/master

Everything else the page reads live: the [access] push line out of the tree, and who you are out of your session. That means it also stays correct later. Sign in, or get added to the list, and the page stops telling you to open a proposal and says you may push now.

The mockup's e91b7d · 2m ago line is not drawn. A time carried in a url is a number the reader supplied to themselves, and they already know when they pushed.

repocfg.List and repocfg.Who were moved out of the hook so both the terminal and the page render the config line from one function. That was the point of copying the wording in the first place.

A run's footer named the wrong ref, and a rev that is not here drew a page

run.html puts the triggering ref in the footer, refs/proposals/46. Forge put the default branch there, which for a proposal build is the one ref the run had nothing to do with.

Separately, /john/johnbot/run/8 answered 200 with "nothing has been built at this commit". 8 is a valid rev name and not a commit in the repository, so the page was drawn for something that does not exist. A rev that does not resolve is a 404 now. A real commit with no runs still gets the page, because that is a true state and a reader may have come looking for it.

A collapsed diff had no expand link, only a way off the page

Chapter 24 on the repository log: "Diffs over a threshold collapse with a size label and an expand link." The size label was there. The link said view commit and went to the commit page.

repo-log.html reads 3d · large diff collapsed · expand. Forge read 14 files · +302 -288 · view commit. Both the word and the destination were wrong: expand means show it here, and the sha at the front of the row is already the link to the commit page.

The link is ?expand=<sha>#<short> now. It reopens the log with that one commit open and jumps to it, keeps ?from= so a reader on the second page stays there, and every other row stays collapsed, which is the point of collapsing.

The diff itself comes from the cache the log already wrote. The first version read only the cache, which meant the link silently did nothing whenever the cache was cold. The end to end test caught it, because the test harness does not set the caches. It falls back to reading the one commit now.

The releases page crossed its budget again, and the fix was one commit hash

Adding the open thread count to every repository page cost about 1.5ms, which put releases at 10.8ms against 10. Measured rather than guessed:

Releases  2.7ms   ReadNotes  4.4ms   OpenCount  1.4ms

ReadNotes was walking the notes tree through the object pool on every view, reading about forty objects to draw twenty bodies. Notes live under one ref, and that ref is a commit, and a commit is immutable. So the whole object-to-body map is read once and cached under the notes commit.

The cold pass costs more, because it now reads every note rather than the twenty on screen: 11.5ms once, then 1.4ms for every view until somebody pushes a note.

Releases  2.6ms   ReadNotes  1.4ms   OpenCount  1.4ms

The releases row is 5.9ms. Every row of chapter 25 passes.

The tree walk also lost the filter that kept it to wanted notes, and gained the rule it should have had from the start: a tree entry whose accumulated path is a full object id is a note, and a shorter one is a fanout directory.

Every page now has to render the same with no cache at all

Last pass the expand link silently did nothing on a cold cache, and the only reason that was caught is that the end to end harness happened not to set the caches. That is luck, not a test.

So the harness sets them now, and every existing test runs the warm path, which is the one production takes and the only one where a wrong cached answer can appear. Then one test takes them away and fetches thirteen pages twice: the profile, the log, the log with a diff expanded, the file tree at two paths, a file, a commit, a compare, the threads list, runs, releases, config and a search. Every pair must render identically, relative times aside.

An empty cache is not a cold cache. The first version deleted the cache directory, which proved nothing: cache.Disk keeps a hot map in memory, and even without it the log writes every diff it reads before anything asks for one back. The honest test is a cache that retains nothing, which is what a full disk is, so the second pass sets them to nil.

And a test needs something to find. The second version still passed with the fix removed, because the repository had two small commits and nothing to collapse, so the expand url changed nothing either way. The repository has a seven file commit now. With the fallback taken out the test fails by name and by url.

That is three versions of one test, two of which proved nothing. Writing the test is the easy half.

The invariant it holds is the one chapter 10 item 6 states: caches are derived, discardable, and rebuildable. If deleting them changes what a reader sees, one of them is not a cache.

Four of the five limits were settings that did nothing

[limits] in the server config has five keys. One of them, signup_per_hour_per_ip, is enforced. The other four are values an operator can set and forge never reads: max_blob_mb, max_push_mb, max_open_proposals, artifact_retain_days.

That is worse than not having them. A config key the file accepts and the server ignores is a promise the operator has no way to check.

max_blob_mb is enforced now, because chapter 20.2 is the one that cannot wait: "Turning LFS off is not enough on its own. Without a limit, a user commits a 4 GB video directly into git. That is worse than LFS, because it is in the history permanently and every clone pays for it forever."

The check runs in pre-receive over the pushed range and not the whole repository, which inside a hook is exactly rev-list --objects <new> --not --all, since the ref has not moved yet so --all still holds the old tips. Two processes, and only when a limit is set. The ids go through one cat-file --batch-check.

rev-list --objects prints the path beside each blob, which is the whole point: chapter 20.2 says "The message must name the file and its size. A rejection that says only 'push too large' sends the user hunting." So the rejection reads

demo.mov is 2.1mb. the limit is 1mb.

large files belong in object storage, with a url or a checksum in the
repository. the build fetches them.

The second half is chapter 20.4, which says to put it in the rejection rather than leave the user with a refusal and no direction. Blobs over the limit are sorted largest first, and a push with several says how many, so a reader fixes the worst one first instead of pushing five more times.

The test pushes a two megabyte file against a one megabyte limit and asserts the name, the limit, the object storage line, and that the small file in the same commit is not blamed. Removing the check fails it.

Still unenforced, and recorded rather than half done: max_push_mb, max_open_proposals and artifact_retain_days. max_open_proposals is BUILD.md's "rate limit proposal refs per key per repo, rule 5 is an open door", and it belongs with the other trap on the same list, expiring unreferenced proposal refs.

The open door had no doorstop

Rule 5 says anyone authenticated can propose. Chapter 27 opens with "Rule 5 is an open door and must be defended without closing it", and names the defense: "Cap open proposals per account per repository at a small number, ten is plenty."

Appendix D's pre-receive pseudocode has the check on the line after the one forge already had:

if not may_propose(user, config): reject("...")
if open_proposals(user, repo) >= limits.max_open_proposals: reject("...")

Forge had the first and not the second. max_open_proposals was one of the four [limits] keys nothing read.

An open proposal is an open thread that carries a ref, which is what the threads list already means by the word, so the count comes from the same place. Three things the test pins down, because each is a way to get this wrong:

  • Per account. Mark at the cap does not stop john proposing.
  • Only on opening. A force-push to refs/proposals/2 is revising something already open, not opening another, and chapter 12 makes that the normal way to update a proposal. Capping it would break the mechanism it is protecting.
  • The message says what to do. "you have 2 proposals open on john/johnbot. the limit is 2. land or close one, then push this again."

Removing the check fails the test.

Still open on the same list

Chapter 26 names two more, both about a repository growing without bound, and both are config keys that already exist and do nothing:

  • Proposal refs. "Expire proposals with no activity for [proposals] expire_days, default 180. Delete the ref, retain the thread. The thread is small; the ref pins commits." repocfg.Proposals.ExpireDays is parsed and defaulted and never read.
  • Proposal revisions. "Keep the most recent five and the ones with anchored comments. Delete the rest on the same expiry schedule."

They belong together, in Server.Sweep, which already runs on a schedule and already drops expired tokens and old events. The anchored comment rule is the part that needs care: chapter 43.4 keeps a comment's line visible after the code moves, and it does that through the blob hash recorded beside it, so a revision holding one of those blobs cannot be deleted.

A proposal ref pins commits forever, and expire_days did nothing

Chapter 26 names three places a repository grows that other forges do not have. The first: "Proposal refs. Anyone may create them, so they accumulate. Expire proposals with no activity for [proposals] expire_days, default 180. Delete the ref, retain the thread. The thread is small; the ref pins commits."

repocfg.Proposals.ExpireDays was parsed, defaulted to 180, and never read.

The sweep already runs on a schedule and already drops expired tokens, old events and trash past its window, so this went beside them. Per repository, per proposal ref: the thread's own last activity decides, since a thread and its proposal are one object and the thread is where activity lands. A ref with no thread behind it is dated by the commit it points at, read through the object pool rather than a process.

A window of zero is expiry switched off, not expiry of everything. That is the kind of default that deletes a whole forge on a config typo, so the test says it out loud.

The commits are not gone, and the test asserts that too. update-ref -d unpins them and git gc collects them later, which is the same order chapter 26 puts them in: "Run git gc per repository on a schedule, not on push. Repack after bulk ref deletion, or the pack files retain everything you just deleted."

repo.Walk came out of this. The reindex command had its own copy of the walk that skips the trash directory, and that copy is where the bug two passes ago lived, so there is one of them now and both callers use it.

Still open, the other two thirds of chapter 26. Revisions retain the old tip on every force-push under refs/revisions/<n>/<k>, and the rule is to keep the most recent five and the ones with anchored comments. The retention half is easy; the anchored half is not, because chapter 43.4 keeps a comment's line visible through the blob hash recorded beside it, so a revision that holds the only copy of an anchored blob cannot be deleted without breaking a comment that is still on the page. git gc per repository on a schedule is also not run.

The field that says which revision a comment belongs to was never filled in

Chapter 26's second growth point: "Proposal revisions. Each force-push retains the old tip under refs/revisions/<n>/<k>. Keep the most recent five and the ones with anchored comments."

The first half is arithmetic. The second half needs to know which revision a comment is anchored to, and thread.Comment has had a Revision field, written into the note format and parsed back out of it, since threads were built. Nothing ever set it. Every comment in every repository says revision 0.

So the retention rule had no input, which is presumably why the pruning was never written.

A comment now records the revision it was written against. That number is the one the content on screen will take when the next force-push retains it, which is the count of existing revision refs plus one. proposal.CurrentRevision reads it from the ref files, so a comment costs no process to number.

nextRevision in the hook was forking for-each-ref for the same count. It calls CurrentRevision too now, so a proposal update starts one process fewer.

A comment that will not say pins everything. Every comment written before this pass says revision 0, and chapter 43.4 keeps a comment's line visible through the blob recorded beside it, so deleting the revision that holds that blob breaks a comment still on the page. When an anchored comment cannot name its revision, nothing is pruned for that proposal at all. It is the conservative answer and it un-sticks itself as comments are written.

An expired proposal keeps none. The sweep prunes with a keep of five normally and zero for a proposal whose ref it just expired, because the reason the ref went is the reason the revisions should go with it. Anchored comments still hold what they need, since the thread outlives both.

Eight revisions, five newest kept, one comment anchored to revision two: one and three are deleted and two survives. Removing the anchored check prunes it and the test says so by number.

The last third of chapter 26, and the sentence that made it urgent

"Run git gc per repository on a schedule, not on push. Repack after bulk ref deletion, or the pack files retain everything you just deleted."

The second sentence is about what the last two passes built. Expiring a proposal ref and pruning retained revisions frees nothing on their own: the objects stay in the pack, so a repository that grew without bound still grows without bound, only with a shorter ref list.

The sweep collects each repository now, and which form it runs is decided by whether it deleted anything there:

  • It deleted refs: a full git gc, which repacks. That is chapter 26's "repack after bulk ref deletion", run the moment the deletion happens rather than left to a threshold that may never trip.
  • It deleted nothing: git gc --auto, which is git's own scheduled form. It costs one process and returns immediately unless git thinks there is work.

No --prune=now, deliberately. Git's default two week grace exists because a push in flight writes objects before it writes the ref that reaches them, and pruning aggressively deletes them out from under it. So a proposal ref expired an hour ago is unpinned now and collected a fortnight later, which is the same order chapter 26 states and the reason the expiry test asserts the commit is still readable straight afterwards.

The test packs a loose repository, checks a pack file appears where there was none, and checks the commit a ref still reaches survived it. A repack that loses reachable objects is the one way this can be badly wrong, so it is asserted rather than assumed. Taking the gc out fails it.

Chapter 26 is complete. All three growth points it names are handled: proposal refs expire, revisions are pruned to five plus the anchored ones, and note trees are left alone on purpose, because 26 says the durability of discussion is the product.

The whole push has a limit too, and it costs nothing extra to check

max_push_mb was the third of the four [limits] keys nothing read. Chapter 20.2 sets it beside max_blob_mb, and config.go already carried the reason it is loose: "the push most likely to hit it is somebody's first import."

Both limits ask about the same objects, so they are one pass of git. The pre receive check reads the pushed range once, sums every new object, and picks the largest blobs out of the same output. Two processes for both limits, and none at all when neither is set.

The two rejections are deliberately different, because the fixes are:

demo.mov is 2.1mb. the limit is 1mb.
large files belong in object storage...

this push adds 1.4mb of objects. the limit is 1mb.
push it in parts, or ask whoever runs this forge to raise [limits] max_push_mb.

One is the user's mistake and chapter 20.4 says where the file belongs. The other may be a perfectly good repository meeting a policy, so it names the setting.

It says "adds ... of objects", not "is". The number is the uncompressed size of what the push adds, which is what forge can measure at pre-receive and what decides how much disk it takes. Saying "this push is 1.4mb" of a transfer that was four hundred kilobytes on the wire would be a number that does not match anything the user can see.

The test pushes three files, each under the blob limit and together over the push limit, and asserts the rejection does not name a file, because naming one would mean the wrong limit fired.

What is left of the four

artifact_retain_days is the last one, and it cannot be enforced because there is nothing to retain. [paths] artifacts is configured and the directory is created, and no code writes to it or reads from it. Chapter 22 is unbuilt: a release has a body in refs/notes/releases, which works, and attached files, which do not exist.

That is the honest state. Chapter 22.4's rule is already written down for when they do: build artifacts expire, release artifacts do not, because a published download that disappears breaks other people's installers.

Release artifacts exist now, and chapter 22 is built

[paths] artifacts was configured, the directory was created at init, and no code ever wrote to it or read from it. A release had a body and no files.

The upload is chapter 22.5 exactly. "A run can attach its output to a release. The job token from chapter 15 carries the permission, scoped to one repository and one job." POST /runner/artifact takes that token, and it is the job's own token and not the runner's long-lived one, which is the difference that makes "one job" true. The poll already issued it and labelled it job <id>; that label is what binds the token to the job, and the test proves a plain git token for the same repository is refused.

The tag has to be a tag in the repository. Without that check an upload creates a directory nobody can ever reach, which is a disk leak with no page to show it.

The build gets what it needs to speak the protocol. BAREREPO_URL, BAREREPO_REPO, BAREREPO_JOB and BAREREPO_JOB_TOKEN are in the environment of [build] command, so a build attaches a file with curl and forge invents no new syntax to describe artifacts. The runner setup page lists the endpoint beside the other four and shows the command, per chapter 24's rule that the page says what the protocol is so anyone can write their own runner.

The write is a rename. A half finished upload is a dotfile ending in .part, and List skips it, so a reader never sees a truncated binary. A rerun replaces a file rather than appending to a list.

The download is an attachment and never a page. application/octet-stream, Content-Disposition: attachment, nosniff and a sandbox policy, which is what chapter 42.3 asks of any bytes a stranger uploaded. Read access is checked, so a private repository's binaries are not public.

A deleted repository takes its files. They are not in git, so nothing else would have.

artifact_retain_days stays a setting that does nothing, correctly

Chapter 22.4: "Build artifacts expire. Release artifacts do not, because a published download that disappears breaks other people's installers."

Everything built here is a release artifact, so nothing expires and the key has nothing to act on. Build artifacts, the kind that expire, are files a run keeps without attaching them to a release, and no mockup shows them: chapter 24's run detail entry is "status, exit code, what triggered it, and the complete log as plain text", with no artifact row. Inventing that surface to give the key a job would be building a page the plans do not describe.

So it stays unused on purpose, and this is the note that says why rather than leaving the next reader to find a dead key and guess.

The same bug as last time, in the new feature

Auditing the artifact work found the mistake the search index made two passes ago, in the same shape: what changes this state without going through the path I built?

Attached files live at <artifacts>/<owner>/<name>/<tag>/. A rename changes the name and a transfer changes the owner, and serveRepoMove moved the directory, the database row and the search index, and left every attached file behind. The releases page would show none, and the bytes would sit on disk with no page to reach them and no delete to collect them, because a delete only removes the path the repository has now.

artifact.Move runs beside MoveDocs now. The test renames, checks the file is still listed and still downloads, then transfers to another account and checks again, because the two halves of the path move separately.

A copy does not take them, and that is right. Chapter 21.1 lists what copying carries: branches, tags, all history, threads and notes. Attached binaries are not on that list, and 22.3 already says a mirror does not take them either.

Two uploads of one name could write into each other

Put wrote to .<name>.part and renamed. Two jobs attaching the same file name to the same tag at the same time share that path, so one truncates the other and the rename publishes a mixture. os.CreateTemp gives each upload its own part file now.

The releases page did twenty directory reads to draw nothing

Listing files per release meant a ReadDir per row, twenty of them, on directories that do not exist for a repository with nothing attached, which is almost all of them. The row sat between 5.9 and 12.2ms against a 10ms budget and crossed depending on what else the machine was doing.

The whole blob store for a repository is read in one pass now: one ReadDir of the repository's directory, and one more only for a tag that actually has files. A repository with nothing attached costs one failed stat.

That row has been near its limit for several passes and every small addition tipped it. This is the difference between nudging it under and giving it room.

A rename detached every runner, and six other things

The checklist from the last pass turned into a test, and the test found six more holes in the same wall.

DB.Move updated two tables, repos and redirects. Seven others key on the repository path and none of them moved:

  • tokens.scope. A runner token is scoped to john/johnbot. After a rename the scope still says the old path and every job is queued under the new one, so the check in jobRunner never matches again. The machine polls forever and builds nothing, and the runners page is empty. Renaming a repository silently detached every runner attached to it.
  • runners.repo, so the page could not list them either.
  • jobs.repo, so anything already queued was orphaned.
  • events.repo and participation.repo, so a transfer left the inbox entries with the old owner and gave the new one nothing.
  • webhooks.repo, so a failing hook's count reset to zero and the config page said it had never delivered.
  • search_docs.repo, which was moved separately by the handler, one more place to forget.

All seven move inside the same transaction now, and the handler's separate call is gone. The test writes one row into every table, renames and transfers in one move, and asserts each one followed. Before the fix it fails seven times.

The budget was measuring a repository forge does not keep

The releases row had been flaking between 5.9ms and 12.2ms against its ten, and the last pass's fix was not the reason it passed.

The cause was reading a hundred loose tag ref files. Under any disk load that doubles. The search row, one SQLite query, barely moved in the same runs, which is what said it was I/O and not the machine.

Forge gcs every repository in the sweep now, and gc packs refs. So a repository forge has been hosting for an hour reads its refs out of one file, and the benchmark was measuring one that had never been swept.

The test packs refs before measuring, because that is the state forge maintains, and the whole table changed:

file tree             0.6ms   was 1.9
threads list          1.0ms   was 3.5
one thread            0.7ms   was 1.9
releases              0.9ms   was 5.9

Nothing is near its limit now. The honest caveat: a repository between a large tag push and the next sweep does have loose refs and is slower, for up to an hour. That is a real state and it is bounded by the sweep, which is the reason the sweep exists.

Worth noticing that garbage collection turned out to be load bearing for page speed and not only for disk.

The delete had the same six holes as the rename

The move test made the shape obvious, so the same test was written for the other end of a repository's life. DB.Forget dropped one row, the ownership row in repos, and left six tables pointing at a repository that no longer exists.

What that looks like to a user:

  • The keys page lists a runner token for a repository that is gone. Its detail line names john/johnbot, which 404s. There is no way to tell from the page that the token is now worthless.
  • A machine stays attached to a repository with no page, polling forever.
  • A queued job waits for a build that can never run.
  • The inbox keeps its lines, each linking to a 404, which is the one thing chapter 24's own rule about names says not to do.
  • A failing webhook's counter survives, so a repository created later with the same name inherits somebody else's failure count. That one is not just untidy, it is wrong.
  • Search still finds it. The handler was dropping the index separately, so this one was covered by accident rather than by the store.

All seven deletes are one transaction in Forget now, and the handler's separate index call is gone, the same consolidation the move got. ForgetDocs and MoveDocs are both deleted: a caller that has to remember a second call is a caller that will forget it.

A delete is safe to be this total because there is no restore. Chapter 44.4 keeps the git data in trash for thirty days, and that is the recovery path. None of these rows are recoverable state: a token is a secret nobody can read back, a runner reattaches with one line, a queued job reruns on the next push.

The test is the delete half of the move test, sharing the fixture that fills every table. Both fail loudly when the code is taken out.

The old name was freed by the one door that does not go through the transport

Chapter 21.2 states it in bold: "The old name is never freed. This contradicts the instinct to recycle unused names, and it is deliberate."

A push to a renamed repository already followed the redirect, because the transport looks the redirect up before it considers creating anything. The form at /new does not go through the transport. It called repo.Create directly, and repo.Create only asks whether a directory is there. The directory moved, so the name looked free.

Take john/johnbot to john/ircbot, then make john/johnbot again on the form, and there are now two truths: a real repository at the old path, and a redirect row saying that path is somewhere else. Everything that points at the old name is wrong, and chapter 21.2's reason for the rule is the worse half of it, that the new repository inherits every mention of the old one.

And repo.InTrash had never been called. Its own comment says "reports a name still in its window, where a push fails rather than creates. Appendix D." Nothing called it, from either door. So a repository deleted a minute ago handed its name straight back out while thirty days of its data sat in the trash under that name.

transport.Claimed answers both questions in one place and says which it is:

john/johnbot is now john/ircbot. the old name is kept forever, so nothing
that points at it breaks.

john/johnbot was deleted. its name is held for 30 days, then it is free.

Both doors ask it now, mayCreate for a push and serveNewRepo for the form. The test takes both doors for both cases, and fails on both without it.

The shape, again. Two ways in, one of them checked. It is the same mistake as the delete that dropped one row of seven and the rename that moved two tables of nine. The question that keeps finding it: what is the other way this happens?

A reply pushed from a clone told nobody

The two doors again, and this time the one that was silent is the one the whole design is about.

Chapter 3's claim is that discussion is git notes in your clone. Chapter 19.2 says participation is subscription: "Anything in a thread you opened or replied to. There is no watch button and no subscribe button."

recordPushes skipped refs/notes/ entirely. So a reply written on the web recorded thread.replied and subscribed its author, and the same reply pushed from a clone recorded nothing and subscribed nobody. The owner of the repository never heard it. The person who wrote it never heard the answer.

Every event kind chapter 19.1 lists for threads, thread.opened, thread.replied and thread.closed, could only ever be produced by the web. The door the book calls the point produced none of them.

Post-receive reads the thread's meta at the old commit and at the new one, which is two pooled object reads and no process, and decides from the pair:

  • no meta before it, so the ref is new: opened
  • open before and not open after: closed
  • otherwise: replied

Reading both sides is what stops a closed thread reporting itself closed again on every later push.

The test opens a thread on the web and replies to it by pushing a note, then reads two inboxes: the owner's, which must have the reply, and the replier's, which must now carry the thread he joined by pushing to it. Without the fix both are empty.

A failed build told nobody either

Chapter 19.1 lists nine event kinds. Eight of them had a writer. run.failed had a constant, a line in eventLine to render it, and nothing anywhere that recorded one.

The chapter is not ambiguous about whether it should exist. It names run.failed in the list and then says, on the next line, "run.succeeded is not an event. A green build is not news." The whole sentence is there to draw the line on one side of which a red build sits.

And chapter 19 opens by saying why any of this exists: "a proposal arrives and the owner finds out by chance. A forge nobody hears from is broken."

So serveRunnerDone records one when the exit code is not zero, and records nothing when it is zero, which the second test asserts, because a feed that reports success is a feed people stop reading.

The number is what makes it reach the right person. The event carries the proposal number when the ref is a proposal ref, so it lands in the inbox of whoever opened that proposal, per chapter 19.2's "anything on a proposal you opened". Without it a failure would only reach the repository's owner, and the person whose change broke would be the last to know.

The actor is the machine that ran it. A build has no human author, and the runner name is the true answer to who is reporting.

The sweep that found it

Listing every event kind against the code that writes it took one command and found the one gap:

for k in ProposalOpened ... ; do grep -rn "Kind: *store.$k" ...; done

Two passes ago the same shape found the thread events, which only the web produced. It is worth doing for any set the book enumerates: the book lists nine, the code should write nine, and anything with a name and no writer is a promise nobody keeps.

Chapter 42, swept section by section

"Everything in this chapter is a security requirement. None of it is optional." So each section was checked against the code rather than assumed.

42.1 markdown and 42.4 file rendering are done and were already right: an allowlist of elements and attributes, on* and style stripped explicitly, script and its siblings dropped with their contents, a null byte in the first 8000 bytes marks a file binary, a megabyte caps rendering, and both cases draw a notice and a download link.

42.3 file content and 42.6 ref names likewise: attachment headers with nosniff and a sandbox policy, ValidRef before any ref reaches an argument, and argument arrays everywhere so a ref named --upload-pack=evil is a name.

Two sections were not done.

A blocked image was blocked silently

42.2 ends: "Proxy remote images through the server or block them. A remote image in a comment leaks the reader's IP address to whoever posted it. Blocking is simpler and honest; say so in the UI."

Forge blocked it and said nothing. The source attribute was dropped and the <img> was written anyway, so the reader got a broken image icon and no reason. Worse, the comment above the code read "it is blocked and the page says so", which was not true, and a comment that describes behaviour the code does not have is worse than no comment.

A blocked image is now replaced, not emptied:

remote image blocked, it would tell its host who read this

The reason is in the sentence because the reader is the person it protects, and a notice that only says "blocked" reads like a bug in forge rather than a choice made for them.

The test chapter 42.5 asks for by name did not exist

42.5: "Derive this list from the route table in code rather than copying it. A route added without a matching reservation is a route an account can shadow, and that is a bug the test suite should catch rather than a list a person must remember."

names.go said, in a comment, that TestReservedCoversRoutes kept the list in step. There is no such test and there never was. A comment naming a test that does not exist is the same failure as the image comment on the same day.

The test now parses web.go, finds every comparison against r.URL.Path, takes the first path segment of each literal, and requires it to be reserved. That is derived from the route table, because the switch is the route table.

It found /signout unreserved on its first run. An account named signout could be registered, and GET /signout would then draw that account's profile while POST /signout ended your session. Reserved now.

Two other things the sweep is worth repeating for: it reads the switch, so a route added tomorrow is checked tomorrow, and it fails loudly if it finds fewer than eight routes, which is how it says it has stopped reading the right thing.

A webhook could ask for an event that does not exist

Chapter 23.1 calls a webhook the escape hatch, "the thing that makes it acceptable to refuse every integration request forever". An earlier pass wrote down what that means: an escape hatch present in the source and absent at run time is worse than one never started, because the config file accepts the lines.

A misspelt event name is the same failure in a smaller box. Write

events = ["push", "proposal.open"]

and the file parses, the hook is stored, and the second name never matches anything. The config page said "delivered 2m ago" because the first name worked, and nothing anywhere said the second one was a typo.

run.succeeded is the sharp case. It is a name a person will reach for, and chapter 19.1 mentions it exactly once, to say it is not an event. A hook asking for green builds waits forever and looks healthy while it does.

The config page names them now, and it is the same page that already reports a failing hook, because it is the only report a hook has:

run.succeeded is not an event forge sends, so it never fires

Chapter 19.1's nine kinds are a list in the store now, with the check that reads it, and a test asserts every kind in the list is known by the check. A list and a lookup that can disagree is a bug waiting for a tenth event.

Why report rather than reject the push. A pre-receive rejection over a typo in a hook would refuse code because of a line about notifications, and chapter 14 is clear that a bad config file must not lock anyone out. The page is where a webhook's state already lives.

The env block was the one place a skipped step was silent

Chapter 15A rests on one rule, and states it in bold: "a skipped step is never silent. A build that reports success while having quietly run half of what was asked is worse than no build at all."

Every decline the chapter lists was implemented and reported: a step using an action forge does not run, a step or job behind an if:, a shell forge cannot start, an expression outside the substitution list. Both enumerated lists were complete too, the twelve environment variables GitHub sets and the eleven expressions forge fills.

env: was the hole. It was written before any of that and never revisited.

An expression in an env value was neither filled nor reported. The chapter explains why substitution exists at all: "${{ reaches sh as a bad substitution and fails the step outright." In a run: line that is true and loud. In an env: value it is not, because the value is shell quoted, so

env:
  TOKEN: ${{ secrets.NPM_TOKEN }}

exported TOKEN holding the literal text ${{ secrets.NPM_TOKEN }} and the build carried on. Nothing failed and nothing was said. secrets is the case chapter 15A names first among the expressions forge declines, and it is where every real workflow puts one.

A value that is not a scalar was exported empty. scalar returns "" for a list or a map, so an env block forge could not read became export LIST=''. An empty string is a value, and exporting one is a claim about what the workflow asked for. It is left unset and named now.

Both go through the same substitute-and-report path a run: step uses, so the three env scopes, workflow, job and step, all report against their own name.

Why this one hid. Every decline the chapter enumerates was there, so a sweep of the list came back clean. The hole was not a missing item on the list, it was one code path that never asked the list anything.

The paste block told a private repository to make itself public

My own bug, from the pass that added a description to the new repository form.

The empty page's block already had a Public flag, and I reused it to choose the visibility line in the new heredoc. That flag does not mean what its name says. It is res.Config.Public() && !wantsPublic, which exists to decide whether to offer the make-it-public line at all, and for an empty repository it is always false, because an empty repository has no commits and therefore no .barerepo/config to be public in.

So the heredoc always wrote visibility = "public". Choose private on the form, give it a description, and the block forge hands you publishes it. Chapter 11 is the reason that is the worst possible direction for the mistake: "accidentally publishing code is not recoverable, and accidentally hiding it is one line in a file."

The page carries the word the form was told now, rather than inferring it from a flag that means something else.

The test suite could not have caught it, and now can. TestTemplatesRender gives each template one set of data, so a branch that data does not reach is never rendered. The repo-empty case has no description, so the heredoc arm was never executed and the missing Visibility field never errored. A second end to end case covers the described private repository, which is the arm that was wrong.

Worth writing down as a shape: a template with a branch has states, and rendering one state is not rendering the template.

The sweep that found it

Every command forge prints in a <pre class="box">, read against what it would actually do. The rest hold up: the clone-and-push copy commands name the same ref set the server side copy moves, the runner lines carry the token, the tag commands are stock git, and the config printf is the mockup's own line.

A test for the class of bug, not the bug

Last pass a template branch nobody rendered hid a field nobody supplied, and a private repository was told to publish itself. The fix was one line. This is the guard.

TestTemplatesRender gives each template one set of data. Go templates resolve a field when they reach it, so a field inside a branch that data does not take is never looked up and never errors. Twenty five templates carry about a hundred and thirty branches between them, so rendering one state each leaves most of them unread.

TestEveryTemplateFieldExistsOnItsData reads the parse tree instead of executing it. It starts at layout, because a page file on its own is only its define blocks, follows {{template "x" .}} wherever the dot is unchanged, and collects every field read against the page's own dot. Range and with bodies are left alone, since their dot is a different type. Each field is then looked up on the fixture's struct by reflection.

It failed on its first honest run. repo-empty asks for .Visibility, added to the handler last pass, and the fixture never gained it. So the very fixture I had just fixed the bug in was still out of step with the page, and the render test was still happy.

Falsified twice: once by the real gap it found, and once by renaming .CopyTo to .CopyToo in the config page, which it names by template and field.

What it does not do. It does not prove a branch produces the right words, only that the words it asks for can be found. The private repository case still needs its own end to end test, and has one. This catches the cheaper half of the problem everywhere rather than the whole problem in one place.

The front door had never been opened by a test

Every test in this tree that needed a signed in reader made a session by calling db.NewSession directly. That is the right shortcut for a test about the keys page or the inbox, and it meant the thing those shortcuts stand in for, signing up and signing in, was the one flow nothing exercised. If it broke, every test would still pass and nobody could use the site.

There is one now, with a real key and real signatures:

  • ssh-keygen -t ed25519 makes a key.
  • The signup form takes the name and the public half and answers with a nonce.
  • printf '%s' '<nonce>' | ssh-keygen -Y sign -f <key> -n barerepo-signup - signs it, which is chapter 31.3's line: one command, stdin to stdout, nothing on disk.
  • The account exists and the reader is signed in.
  • /auth/challenge then /auth/verify take the other door, the one a returning reader uses, with -n barerepo-auth.
  • The session opens /keys, since a session that opens nothing is not a session.

The namespace separation is asserted, because it is the reason it exists. Chapter 10: "The namespace is barerepo-signup, not barerepo-auth. With one namespace a signature captured from a sign-in could be replayed to claim an account with somebody else's key." The test signs a signup nonce with barerepo-auth and requires it to be refused. Setting the two constants equal fails it by name.

One thing the test taught me about the code rather than the other way round. A wrong signature spends the nonce. My first version treated that as a bug and it is not: it stops a signature being guessed at against one challenge, the form comes back with the name and key already filled, and the page says "the nonce is spent, so press create for a new one". The test asserts that sentence now, because without it the next attempt looks like forge is broken.

The runner protocol was documented and never driven

Same question as last pass: what does every test work around? Every test needing a runner called AttachRunner and TakeJob on the database directly. So /runner/attach and /runner/poll were named on the add-a-runner page, listed in appendix C, and exercised by nothing.

That page makes a specific promise, and chapter 24 explains why it matters: "the page says what the program does, and says that the protocol it speaks is the plain HTTP in chapter 15, so anyone who would rather write their own has everything they need to. That is the difference between a required tool and a hidden one."

The test is that anyone. It attaches with a token, checks the runners page lists the machine, pushes a [build] command, polls and receives the job, posts a log chunk, posts the result, and reads the run page for the machine name, the log and the status. Nothing in it imports forge's own runner.

It passed first time, which is the honest outcome and still worth having: the claim on that page is checked now rather than asserted. Breaking the log endpoint fails it on the run page, which is where a reader would notice.

What is still worked around. Every push in this suite goes over https with a token in the url, because that is what a test can do without an sshd. So sshx.Serve and forge ssh, the entry point authorized_keys forces and the one chapter 41.3 calls attacker-controlled, are covered only by unit tests of Parse and Write. That is the next hole of this kind, and it is a bigger one, since ssh is the transport the design is built around.

ssh, the transport the design is built on, had never carried a byte in a test

Every push in this suite goes over https with a token in the url, because that is what a test can do without an sshd. So sshx.Serve, the function that takes SSH_ORIGINAL_COMMAND and hands git the connection, and forge ssh, the entry point authorized_keys forces, were covered by unit tests of Parse and Write and nothing else. Chapter 10 makes an ssh key the identity and every page prints an ssh clone url.

A test needs no daemon. GIT_SSH_COMMAND points git at a five line script that does what authorized_keys does: put the command in SSH_ORIGINAL_COMMAND and exec forge ssh --account john. Then a real git clone and a real git push go through the real path, and the log page is read to prove the push landed.

The script taught me something about the shape of the connection. My first one took the first argument as the host and the rest as the command, and git refused it: git sends -o SendEnv=GIT_PROTOCOL ahead of the host for protocol v2. sshd puts only the last argument in SSH_ORIGINAL_COMMAND, so the script does too.

Chapter 41.3 is asserted directly. Six commands are pushed at the entry point: nothing, sh, a git verb with ; touch after it, one with && whoami, scp -t, and a path escaping the root. Each must be refused, none may crash, and each must say who refused it. /tmp/forge-owned is checked afterwards, since the honest question is not whether an error was printed but whether a shell ran.

Three of my falsifications this pass did nothing at all

Worth writing down because it is a failure of method, not of code.

To check that a test can fail I edit the code, run the test, and put the code back. Three times this pass the edit silently matched nothing, because the pattern I searched for had a leading tab the source did not, and str.replace reports nothing when it replaces nothing. The test passed, I read that as "the test does not discriminate", and I was reading an unmodified binary.

I caught it by running the command by hand and seeing the refusal message that the edit should have removed.

Every falsification asserts the edit applied now. assert s.count(old) == 1 before writing the file, so a pattern that does not match stops the check instead of quietly passing it. A falsification that cannot fail is worth less than no falsification, because it is believed.

Going back over the guards, with an assertion this time

Last pass three falsifications silently edited nothing and I read their passes as information. So this pass went back over the guards whose falsification had never been proven, with a helper that refuses to run unless its edit matched, and that says plainly whether the test caught the break.

Twelve guards checked. Ten caught their break. Two did not, for different reasons, and the difference is the interesting part.

One was a bad test. TestOnlyACodeSearchResultDrawsABox builds searchRow values by hand and renders the template, so it proves the template honours HasText and nothing at all about the handler that sets it. Handing a thread the box search.html keeps for a source line is a handler decision, and no test touched it. There is an end to end one now: a query that matches a thread and no file, and the page must hold no <pre> at all. It catches the break.

One was a bad mutation. The cold cache test compares a page rendered with the caches on against the same page with them gone. Deleting a cache write cannot change that, since output with no cache is the property under test. The mutation did not violate the property, so the pass told me nothing about the test. The same test does catch a real break, which is a cache read with no read behind it, and that was falsified when it was written.

So a falsification says something only when the mutation actually violates the property. A mutation that makes the code slower, or uglier, or differently spelled is not a falsification, and reading its pass as reassurance is the same error as reading an unapplied edit.

And one real gap, found by falsifying rather than by reading

Dropping refs/notes/* from the copy's fetch refspec did not fail the copy test. Chapter 21.1 says what a copy carries: "Branches, tags, all history, threads and notes." The test asserted the branches came and the proposal refs did not, and never looked for the discussion. A copy that silently lost every thread would have passed.

It now checks the notes ref arrived and that the copy's thread list is not empty, and the mutation fails it.

Chapter 40.1 lists nine things a mirror takes, and the test checked seven

Same lens as the copy gap: a sentence that enumerates is a checklist, and a test for it must read every clause.

"This copies the code, all history, every branch, every tag, every proposal ref, every thread, every comment, the config file and every build result."

TestPortability builds a repository, opens a thread, pushes a proposal, comments on a line, merges, records a build, mirrors it, deletes the original, pushes the mirror to a second forge, and reads all of it back. It covered code, history, proposal refs, threads, comments, the config and build results.

Every branch and every tag were the two it did not. The repository it built had one branch and no tags at all, so the two plural clauses were unread. It now pushes a topic branch and an annotated v1.0.0, and reads both back off the second instance, along with the release body, which chapter 22.1 says clones with the repository and which nothing had checked either.

And two mutations, one of which taught me something. Removing refs/tags/ from pre-receive's access switch did not fail it, and that is correct: the owner of the destination is restoring, and chapter 18 gives the namespace owner every ref in their own repository, so the switch is never reached for a mirror push. Bad mutation, not a bad test.

The mutation that does violate the property is turning that exemption off. Then the mirror push is judged ref by ref and refs/notes/runs is refused with "build results are written by the server", and the test fails. Chapter 18 says this outright, that the exemption is what makes chapter 40.3 true, and now something holds it: taking the exemption away breaks taking your repository somewhere else.

Chapter 18 calls itself the complete matrix, so now it is a table

"The complete matrix. There is nothing else." Seven rows, and until now the only thing holding them was reading the pre-receive switch and agreeing with it.

Seventeen cells, pushed for real by four accounts against one public repository where john owns it and lisa has [access] push:

refs/heads/*          owner yes, push list yes, stranger no
refs/tags/*           the same three
refs/proposals/new    a stranger yes, which is rule 5
refs/proposals/<n>    its author yes, push list yes, another stranger no
refs/notes/threads/*  anyone who may read yes
refs/notes/runs       stranger no, owner yes
refs/meta/*           stranger no, owner yes
everything else       stranger no, owner yes

The last two rows are the owner exemption chapter 18 states beside the table, and they are in the same table because they are the same rule.

The first version of this test passed two cells for the wrong reason. Every case pushed the same commit, so once an allowed case had written a ref, the refused case that followed it got "Everything up-to-date" from git and the hook never ran. Two cells were vacuous and green.

Each case pushes its own commit now, and the loop fails outright on "up-to-date", because a push that sends nothing has judged nothing. That guard matters more than the cells: a test that quietly stops exercising the thing it names is the failure mode this whole run keeps finding.

Falsified three ways, each catching it: dropping [access] push from the read of who may push, letting the final default accept instead of reject, and removing the proposal author check.

The tool I built to check my tests corrupted the code it was checking

The falsification helper edits a file, runs a test, and puts the file back. It put it back by replacing the mutation string with the original string. That is only safe when the mutation string is unique in the file.

One mutation replaced a reject(...) call with return nil. return nil appears in that file many times, so the restore rewrote the first one, which lives in an unrelated function, into a call with variables that do not exist there. The package stopped compiling, and I committed it before the suite told me.

Both the corruption and the deletion were repaired in the next commit, and the suite is green again.

The helper keeps the whole file now and writes it back byte for byte, then asserts the file on disk equals what it read. A restore that pattern matches is the same class of mistake as an edit that pattern matches without checking, which is what this helper existed to prevent. It made both errors on the same day.

Two rules out of it, both cheap:

  • Restore by content, never by pattern. Keep the original bytes.
  • Run the whole suite before committing, not the one test the pass was about. The build failure was in a package the matrix test never touches, and go test ./e2e/ was perfectly happy.

Rule 3 was false again, in the way chapter 10 warned it would be

Rule 3: nothing is stored that is not a git object, except a closed list in chapter 10. The chapter prints that list and then says, of itself:

"This list was four items in an earlier draft and the count was wrong. Redirects and artifacts were added to the design without being added here, which made rule 3 false while it was still being cited. The count is stated plainly because a closed list that quietly grows is worse than an open one."

It has happened again. Forge stores fourteen tables. Six of them map cleanly onto items 1 to 4 and 6. Four do not:

  • runners, a machine that dialed in and what it can build
  • jobs, the queue of work waiting for one
  • webhooks and webhook_cursor, a hook's run of failures and where the sender got to

None of that is derivable from git, so item 6 cannot hold it: item 6 says "rebuildable from git alone", and nothing in a repository records that a laptop attached this morning. Chapter 19.4 makes the same distinction for read state and concludes it "would have to be added to the closed list in chapter 10 as a new category rather than folded into item 6". That is the reasoning followed here.

The book now has a seventh item, work in flight and who is doing it, with the reason it is not item 6 written into it, and the paragraph about the count now says the list has grown quietly twice.

Chapter 28 and BUILD.md both said five tables. Chapter 28 uses that number to describe a backup, so a reader counting tables would have found nine more than the book admitted to. Both now describe the database by what it holds rather than by a number that goes stale on the next migration.

And a test, because the chapter's own complaint is that this keeps happening quietly. TestEveryTableIsOnTheClosedList reads every CREATE TABLE out of the migrations and requires each to name the list item that owns it. A new table with no item fails the build. It checks the other direction too, so a claim about storage that no migration makes is also a failure. Both directions falsified.

Two thresholds the book states and nothing measured

Chapter 25 budgets a page 2kb of javascript. It is in BUILD.md's table beside the timings, which the notes call "build-failing thresholds, not aspirations", and every row of that table was asserted except this one.

The answer is 786 bytes, one file, one script tag, which is rule 4 honoured with room to spare. But nothing said so, and the next person to reach for a helper library would have found out from nobody. The test reads every <script src> out of the templates, adds up what they pull from the embedded static files, and fails over 2kb. Padding keys.js past the line fails it.

Chapter 24 ends with the pages that do not exist. Sixteen of them, and the chapter gives two different reasons: the discovery pages are absent per rule 7, because "a page that displays emptiness to every visitor actively harms adoption", and the rest per chapter 1, because "remove them and the five things a forge does still work".

A list of things that must not exist is as checkable as a list of things that must. Nineteen paths are requested and every one has to answer 404, and the same test reads the thread list for a merge button, which is BUILD.md's own trap: "Do not add a merge button. Every request for one is a request to become GitHub."

Both halves falsified. Wiring /explore to the search handler fails the first, putting a merge button on the thread list fails the second.

That second one is worth keeping precisely because it will never fail by accident. It fails the day somebody decides one small button would be convenient.

The one link on the site that could only fail

Crawling every link on fifteen signed-in pages, a hundred distinct urls, found one that did not answer: /inbox.atom, linked from the inbox page itself, returned 401 to the person looking at their own inbox.

The 401 was correct and the message was helpful, "this feed needs a feed token. make one on your keys page." But inbox.html draws atom · feed token as two links, and the first one could never work for anybody. A link whose only outcome is an error is a link that should not be there, or a handler that should answer.

The handler answers now. A request carrying a feed token is served as before. A request carrying no token at all, from a browser that already has a session, is served to that session's account.

This does not weaken chapter 19.5. The token exists for a reason the chapter states: "A feed reader stores URLs in plain text, so a URL that grants write access is a bad idea." That is about what goes in a url a feed reader keeps. A session cookie is not sent by a feed reader and grants strictly more than the feed already, so refusing it bought nothing and cost the link on the page.

Both halves are asserted: an anonymous request and a wrong token are still 401, and only the session case is new.

The crawl is worth repeating after any template change. Ninety nine of a hundred links were fine, which is the ratio that makes reading them by hand a bad use of a pass and a script a good one.

The crawl is a test now, and it found the bug the hand run had missed

Last pass's link crawl was a shell loop over fifteen pages I chose. It found one broken link. Written as a test that follows links rather than visiting a list, and run against a repository with a proposal on it, it found another straight away.

Every file on a proposal's compare page linked to a 404. The compare page builds a file link as /file/<ref>/<path>, and for master...refs/proposals/1 the ref is refs/proposals/1. The route reads one path element as the ref, so refs became the ref and proposals/1/config.go the path, and nothing was there.

A ref holding a slash cannot be one path element. Rather than teach the route where a ref ends, the link resolves the ref to a commit when it holds a slash, which is unambiguous and also survives the force-push that chapter 12 makes the normal way to update a proposal. A plain branch name still reads as itself.

And the crawler taught me one thing about my own tooling. Its first run reported three comment links as 404 that were fine: a page writes & as &amp;, and a crawler that does not undo that asks for a url nobody wrote. Three of the four failures were mine.

The test crawls from seven roots, follows every internal href it finds, stops at three hundred pages, and fails if it reaches fewer than twenty five, because a crawl that stops early passes for the wrong reason. Falsified twice: reverting the file link and reverting yesterday's inbox feed fix each fail it.

That is the shape worth keeping. A list of pages checks the pages somebody thought of. A crawl checks the ones they did not.

A ref with a slash in it, everywhere it is spent on one path element

The compare page's file link was the first of three. Grepping for every url built from a ref found the rest:

  • the releases page, linking each tag to its tree. A tag may hold a slash, and release/1.0 is a spelling plenty of projects use.
  • the config page, linking .barerepo/config to its raw bytes through the default branch. feature/x is a legal branch name and a common one.

Both go through the same resolve now. And fileRef had to grow: release/1.0 is not a ref path, so reading the ref files cannot find it, and git is asked when the files cannot say. The first version resolved a full ref name only and quietly left the broken url alone, which the crawl caught the moment a slashed tag existed.

The fixture is the reason it was caught. The crawl passed before, because the repository it built had one tag named v1.0.0 and one branch named master. A slash is legal in a ref and it is the thing that breaks a url spending one path element on one, so the fixture has both a release/1.0 tag and a feature/login branch now. Same lesson as the mirror test that had one branch and no tags: a clause about a hard case is unread until the fixture contains one.

Nothing linked to the releases page

The crawl reached the releases page for the first time only after a tab was added for it, which is how the missing tab was found: the page answered when asked, and nothing ever asked.

releases.html draws the tab row as log · files · threads 3 · runs · config, and so does every other mockup. None of the twenty four links to releases. Only index.html, the contact sheet, does, and that is a page of the mockups rather than a page of forge.

So chapter 24 describes a view, appendix C routes it, a mockup draws it, and a reader could only reach it by typing the url. This is a sixth tab, and it is a visible deviation from five mockups, taken deliberately: a page nobody can find is worse than a tab row one item longer, and the word traces to chapter 24 and to the mockup's own title.

The crawl now asserts reachability as well as answers. Those are two properties and the second one hid: removing the tab makes nothing 404, it makes a page disappear. Nine pages must be reached from the front door, and removing the tab fails it by name.

A file name with a space in it broke its own diff

Following the lesson that a hard case is unread until the fixture contains one, the crawl's repository gained a file called a note.md and one called c++.md.

The plus turned out to be my crawler again: html/template writes + as &#43; in a url attribute, and a browser reads it back as +. The crawler now unescapes html entities generally rather than the one entity I had noticed, which is the second time that same shortcut has produced a false failure.

The space was real. Every link to a note.md on the log and the commit page ended in %09, a tab, and answered 404.

The unified diff format is where it comes from. A +++ b/ line normally ends at the name, but when the name holds a space git writes a tab after it, because that tab is the format saying where the name ends. Forge took the whole rest of the line, tab included, as the path.

So a repository with one space in one file name had a broken link on its landing page. The parse cuts at the tab now.

Two of the last three bugs have been the same shape: a value that is usually a plain token, spent somewhere that assumes it is one. A ref with a slash in a path element, and a path with a space in a diff header. Both were invisible until a fixture held the awkward case, and both were on the pages a reader sees first.

git has two spellings for a path, and forge only knew one

café.md went into the crawl's repository and produced this link on the log page:

/john/johnbot/file/<sha>/"a/caf\303\251.md" "b/caf\303\251.md"

Git quotes any path outside ascii in its own output, as a C string with octal escapes. So diff --git "a/café.md" "b/café.md" has no b/ in it to split on, and no +++ b/ prefix to correct it either, since that line reads +++ "b/…. Both parses missed and the whole rest of the line became the path.

The fix is one setting, not one parser. core.quotePath=false makes git write the raw bytes, and forge sets it for every git process through the environment it already builds. That fixes every place forge reads a path at once: the diff header, the file list, ls-tree --name-only and diff --name-only in the search indexer, and rev-list --objects in the blob size check.

Writing an unquoter instead would have fixed the one call site I was looking at and left the other four to be found later, one crawl at a time.

The search index had the same bug and no way to notice. A repository with an accented file name indexed it under git's octal spelling, so the search page offered a link to a path that does not exist. There is a test for that now alongside the crawl, because the crawl only reads links and the index is not one.

Both falsified by flipping the setting back to true.

The awkward names live in one fixture

release/1.0, feature/login, a note.md, c++.md, café.md. Every one of them found something, and each one was cheaper to add than the bug it found was to find any other way. The next awkward name goes there rather than into a test of its own.

The two branches that draw a notice instead of the file

The awkward fixture grew a directory, a nested docs/a note.md, a file with a null byte in it, a file over a megabyte, and a file deleted in the commit after it appeared. The crawl went green, which says every link those states produce answers. It does not say the pages are right, because the crawl reads links and these two branches draw prose.

Chapter 42.4 asks for two refusals: "Detect binary files by looking for a null byte in the first 8000 bytes. Do not render binary content. Show the size and offer download." and "Cap rendered file size. Files above 1 MB show a notice and a download link."

Both were implemented and neither was ever rendered by a test, because no fixture had ever contained such a file. They are asserted now, in both directions: the binary page says binary file and offers a download and does not contain the bytes, the large page says too large to render and does not contain a line of it, and an ordinary file still renders and says neither. That last case matters, since two rules that refuse everything would pass the first two checks.

Setting the sniff length to zero fails it, and raising the cap to a terabyte fails it.

The deletion, the directory and the nested awkward name found nothing. Worth saying: most awkward cases do not find a bug, and they are still cheap enough that adding them is the right call. Five of nine have found something so far.

The landing page kept its diffs shut

Every mockup carries one sr-only sentence saying what its page is for. Forge has the same mechanism, {{.Summary}} in the layout, and every page fills it in. So the two sets of sentences can be read side by side, and one pair disagreed.

The mock says: "Repository log with every commit diff expanded inline. This is the landing page." Forge said: "Repository log, newest first, each commit's diff one click away."

Forge's sentence was the accurate one. Each diff sat in <details name="log">, which is shut until clicked, and the name makes the whole page an accordion, so opening a second diff closes the first. Its own stylesheet said so out loud: "a commit's diff on the log page, shut until asked for. one open at a time".

Chapter 24 says the opposite: "Commits newest first, each with message, author, time, changed files, and the diff already expanded. Diffs over a threshold collapse with a size label and an expand link." The mockup draws it that way too: no <details> anywhere in the file, two diffs open, and only the third commit, a fourteen-file merge, collapsed with large diff collapsed · expand.

The threshold is the answer to the size problem, and it was already built. The disclosure was a second answer to a problem that had one, and it cost the page its reason for existing: "People arrive at a repository to find out what changed... The log answers the first directly." A page of shut drawers does not answer directly.

Now the diff renders inline and the collapsed branch is untouched. The stats line lost its · view commit, which the mock does not draw and which the linked sha beside it already does.

The page got smaller: 19kb against 22kb, because the disclosure markup was pure overhead. It was never a page-weight measure. The bytes were always being sent.

Putting the <details> back fails the test.

The method here is worth keeping. A screen reader sentence is a claim about what a page does, written twice by two different people. Where the two spellings disagree, one of them is a bug. The crawl now also fails any page that renders that heading empty, so the pairs stay comparable.

A typo in the config took everything away

Chapter 14 states it in one sentence: "A malformed config file must not lock anyone out. On parse failure, fall back to the last known good version and print a warning to the pusher's terminal."

Forge did the warning and not the fallback. Load returned Default(), and the defaults are not neutral, they are the safest possible answer to every question:

  • visibility empty reads as private, so a public repository went dark
  • [access] push empty means owner only, so everyone else lost push
  • require_runs empty means nothing is required any more
  • [runners] empty means no labelled runner matches
  • [[webhook]] empty means the hooks stop firing

One unclosed bracket did all of that at once. The warning it printed was honest about it: "its settings are being ignored". The code and the book disagreed and the code said so out loud.

The fallback now walks the file's own history. git log -n 25 -- .barerepo/config newest first, and the first version that parses is the one in force. The walk is bounded so a file that has never parsed cannot cost a walk of the whole history, and the result is cached under the same commit key as the file itself, so a repository with a broken config pays for the walk once.

Two warnings, because there are two moments. The config is read from the tip of the default branch, which during a push is still the old version. So the push that introduces the typo used to be the one push that said nothing. It now parses the incoming file too and says the file will not parse. Every later push names the older commit whose settings are deciding.

And the config page said nothing at all. A reader opened it, saw the broken file rendered as if it were law, and had no way to know. It carries the same sentence now.

Three falsifications: returning Default() again, dropping the incoming-file check, and dropping the page's error each fail the test.

Found on the way, not built

The book's config page shows "history, blame, and raw links" and the mock draws history · blame · raw. Forge draws raw. The file view has the same gap. Next pass.

The history link that was a grey word

Chapter 24 gives the config page "history, blame, and raw links", and the mock draws all three underlined. Forge drew raw. The file view drew <span class="muted">history</span>, which is a word styled to look like a control and wired to nothing. That is worse than leaving it out: it promises and refuses.

There is no history route, and there does not need to be one. Appendix C has no /history and no /blame, and chapter 24 argues against a blame page directly: "Blame is not a separate question." So the two links resolve to pages that already exist:

  • blame on the config page is the file view of .barerepo/config, which draws blame in the gutter on every line, always. The config page renders the file as a pre, so this is the only way to see who wrote which line of the policy.
  • history is the log page restricted to one path: /<owner>/<name>?path=<file>. A query parameter on a route that already exists, not a new route.

The log page was already the right page for this. It draws every diff open, so one file's history is that file's changes, each with its diff, newest first. It says what it is restricted to and links back to the whole log.

The diff cache had to be told. It is keyed by commit sha and holds the whole commit's patch. A path-restricted log produces a different patch under the same sha, so the filtered path skips the cache in both directions. An unfiltered log is unchanged and still reads from it: 14.2ms, 19kb, well inside chapter 25.

The feature found its own bug, in the crawl

doomed.md is deleted by the awkward fixture. Its history page linked the file name back to the file view at the branch tip, where the file is not, and the crawl caught the 404 within a minute of the feature existing.

A file with a history and no present tense is a real state, so the page says so: the name is plain text and reads which is not in master any more. Removing that check fails the crawl.

Four falsifications, and three of them were fixture failures first: the render cases in view_test.go had to grow the new fields before TestEveryTemplateFieldExistsOnItsData would go green. That guard has now paid for itself twice.

Three settings the config parsed and nothing read

.barerepo/config is the whole settings surface, so a key that parses and does nothing is worse than a missing feature: the page shows the file as though it were law. Counting reads of every field in the struct found three at zero.

[repo] default_branch. Chapter 33.6 is a recipe: edit it, commit, push, "the server reads the file on push". HEAD was set once, on the first push into an empty repository, and never looked at the file again. It follows the config now, and says so in the terminal. A branch named but not pushed gets a sentence rather than a HEAD pointing at nothing.

[proposals] require_runs. Chapter 37.4 is a section called "Require builds to pass" and it did nothing at all. A team could read that section, write the line, push it, and believe the default branch was protected.

Forge has no merge button by design, so there is only one place this rule can live: the push that puts a commit on the default branch. The hook reads the run notes for that commit and refuses it if a required name has not passed, naming the ones that have not and pointing at the runs page.

The owner is exempt, on chapter 21.3's precedent for archived. Without that, require_runs = ["build"] with no runner attached locks everyone out of the repository including the person who has to edit the file to undo it, and the file lives on the branch they can no longer push to.

[runners], the third, is still unread. It maps a hostname to the labels that machine will take, and a runner already advertises its own labels when it attaches, so the config side is a second opinion with no stated precedence. Left alone deliberately rather than guessed at.

The feature found a silent bug behind it

The first version of the test failed with the build passing. The run note held {"runner":"uproar.local","exit":0} and no name, because the job lookup behind /runner/done selected every column except name. Every run forge has ever recorded from a finished job has had an empty name, and chapter 15A's matrix is grouped by exactly that field.

Nothing noticed, because until today nothing read the name back.

The flash of unstyled content

Reported while the above was being written, and real. /static/barerepo.css answered with no Cache-Control, no ETag and no Last-Modified, because an embedded file has a zero modtime and http.FileServerFS sends no validator without one. With nothing to revalidate against, a browser refetches the stylesheet on every navigation, and the page paints before it lands.

The url carries the version now, which makes the body under it immutable, so it is served with a year and immutable. An unversioned url is somebody's bookmark and gets a minute. Chapter 25's own principle: cache whatever is a function of an immutable thing.

The 2kb script budget test caught this within a minute, because it read keys.js?v={{.Version}} as a file name. It reads the path now.

The front door answered 404 to anything that asked politely

Found by running curl -sI against the sign-in page while looking at cache headers.

GET  /signin -> 200
HEAD /signin -> 404
HEAD /signup -> 400

curl -I sends HEAD. The router matched /signin on r.Method == http.MethodGet, so a HEAD fell past every named route into the generic one-path-element case and was answered as a profile for an account called signin. /signup has no method guard at all, so a HEAD reached the form branch and was answered as a submission with no form in it.

HEAD is a GET that stops at the headers. Go's own server discards the body for a HEAD response, so routing it like a GET is all that is needed. Link checkers, uptime probes, and the unfurler in every chat client use HEAD. Every one of them was being told the sign-in page does not exist.

The two pages that were wrong are the two a stranger sees first.

One GET that a HEAD must not reach

/runner/poll is left on MethodGet alone, deliberately. It does not read a queue, it takes from one: the handler removes a job and answers with it. A HEAD routed there would take a build and throw it away, and no runner would ever see it. A link checker walking the site would empty the queue.

So the inconsistency is the correct state, and it now has a test that says so. The test asserts a queued job survives a HEAD to the poll, which fails the moment somebody tidies the last r.Method == http.MethodGet away.

That is the point worth keeping: a rule with one exception needs the exception written down as a test, or the next person removes it for consistency.

The cache key was the release number, which never moves

The flash of unstyled content was reported again after the fix, and the report was right: the server on 3999 was a binary built at 08:58, before any of today's work. Its html still asked for /static/barerepo.css with no version and no caching. Nothing was wrong with the fix; nothing was running it. Rebuilt, restarted, and one navigation to a second page now issues no second request for the stylesheet.

But the fix had a trap in it. The url carried ?v={{.Version}}, and Version is a const, 0.1.0. It does not move between builds. So a browser that took the stylesheet under ?v=0.1.0 with max-age=31536000, immutable would keep it for a year, and the next edit to barerepo.css would reach nobody who already had it. That is a worse bug than the one being fixed: the flash is a nuisance, a stylesheet frozen for a year is a broken page nobody can clear.

The url carries a hash of the files now, computed once from the embedded directory at startup. A changed stylesheet is a url no browser has ever seen, so immutable is true rather than hopeful, and a release number nobody remembered to bump cannot pin an old file.

The test asserts the tag is not the version and that every page hands out the same one, since a page with a stale tag pins a stale stylesheet for whoever lands there first.

The lesson is about immutable itself. It is a promise, and a promise keyed on something that does not change is a lie with a one year expiry. Cache on the hash of the thing, which is the same rule chapter 25 already applies to every git object forge caches.

Anyone could write a comment in anyone else's name

The thread page prints these two lines and invites a reader to use them:

git notes --ref=threads/1 append -m "your reply"
git push origin refs/notes/threads/1

Chapter 35.5 prints the same pair. Following them put the reply on the page inside the previous person's comment, over that person's name, because git notes append joins with a blank line and forge separates records with a line of two dashes.

So forge printed instructions that misattributed the words of whoever followed them.

The repair uses the difference, not a guess. In post-receive the old commit is still there, so what a push added is exactly the suffix of each note that was not there before. If that suffix is not already a well-formed record it is wrapped as one, authored by the account the push authenticated as. A record that names its own author is left exactly as pushed, because chapter 40.3 restores a repository by pushing its notes and a restore that renames every author is not a restore.

Pulling that thread found something much worse

Testing the exemption above turned up the real bug. A comment body is written into the note as-is, and the record separator is a line of two dashes. So this, typed into the reply box on the web page by any signed-in user:

looks fine to me

--
author: john
time: 1755000000

I approve this change.

renders as two comments, and the second one is signed john with a timestamp the writer chose. Anyone could put words in anyone's mouth, including the owner approving a proposal. Two comments were written and the page drew three.

A body is content and a separator is framing, and content that can become framing is the same bug as SQL injection with the same shape. A line of only dashes now gets one more dash on the way in and loses it on the way out. Two dashes is the separator and writing one always produces at least three, so a body can no longer end its own record. A reader who types --- still sees ---.

Falsified: dropping the escape lets the forged comment through, and the test counts the comments rather than looking for a name, so it fails on the third comment existing at all.

The same bug again, twice, through the front door

Yesterday's separator escape closed one half of the record format. The other half is the header block, and it had the same shape of hole in two places.

A hidden form field chose the name over a comment. The line comment form carries blob, the hash of the file the comment is anchored to, and the handler read it with strings.TrimSpace and nothing else. TrimSpace does not touch a newline in the middle. So blob=abc\nauthor: john wrote:

author: mark
time: 1755...
anchor: README.md:1
blob: abc
author: john
side: new

and the parser takes the last author: it sees. Posted as mark, signed john, from the ordinary comment form on the ordinary page.

A thread title could claim a header of its own. The title is the first line of the meta blob, so title: x\nmerged: <sha> made a thread claim it had been merged. state: happened to be safe only because Render writes it after the title and the last one wins. Safe by accident is not safe.

Both are fixed at the one place that writes a record, not at the handlers. Every header value goes through oneLine on the way out, so a newline in any field becomes a space and a value can never start a line. Fixing this at the call sites would have meant finding all of them, and the next field added would have to be found again.

The rule this makes explicit. A record is lines of key: value and then a body. Nothing that comes from a person may contain the two things that structure it: a newline in a header, or a line of dashes in a body. Both are now escaped where the record is written. That is one place, and it is the only place either rule needs to live.

Expanded diffs, and the budget nobody was measuring

Reported: the log page is a wall of open diffs. It was, and the report found a second thing behind it. The landing page of this repository weighed 219kb. Chapter 25 budgets it at 30kb, and calls the table "build-failing thresholds, not aspirations".

Chapter 24 asks for both: "the diff already expanded", and "Diffs over a threshold collapse". The threshold that existed was per commit, six files or 160 lines. Twenty commits can each sit under it and still add up to seven times the page budget. Two rules that are each satisfied and together are not.

So the page has a budget of its own now. Diffs open from the newest down until the inline diff content reaches 18kb, and the rest collapse with the same expand link, saying collapsed to keep this page small rather than large diff collapsed, because a reader deserves to know which rule shut it. The commit a reader asked to expand is never shut by this, whatever it costs.

This repository's landing page is 27kb now, three diffs open, fifteen collapsed for the page and four for their own size.

The budget test was measuring a repository nobody has

It builds a thousand files and two hundred commits, and every commit changed one line of one file. Twenty of those are 19kb of page, so the test passed while the real thing was seven times over. A fixture can be large and still be nothing like the thing it stands for.

Each commit now changes three files by ten lines each, which is under the per-commit rule and over the page's. Getting there took two wrong fixtures: the first rewrote whole files, so every commit collapsed on its own; the second picked different files each commit without carrying the earlier ones forward, so every commit reverted the one before it and changed six files instead of three. The fixture has to be right before the measurement means anything, and both wrong versions passed.

Removing the page budget now fails the test by 3kb, and so does collapsing everything, because the same test asserts at least one diff is open. Chapter 24 and chapter 25 hold each other in place.

Collapsed by default, and the book says so now

The expanded log page is reverted. Diffs are shut until asked for, one open at a time, which is what <details name="log"> does and what was there before.

Chapter 24 is amended rather than worked around, because the code and the book must not disagree: it asked for "the diff already expanded", and that is a 219kb page on a real repository against chapter 25's 30kb budget. The mockup draws three commits.

A shut disclosure still sends its bytes, so the page budget is still needed. Past 18kb of diff the page stops sending diffs at all, and those commits get the same expand link the large ones already had, which reloads the page with that one diff in it. Every commit opens; only the first few open without a round trip. No new words on the page.

"collapsed to keep this page small" is removed. It explained forge's own budget to somebody who did not ask, which is a note for whoever wrote it and not for whoever is reading. Nothing in the interface should explain the implementation.

Two more budgets measured against a repository nobody has

The same fixture problem as the log page, in two more rows.

The file tree measured /john/big/files, the root, which holds two entries. The thousand files are in src/. Measured there it is 234kb against a 15kb budget. A directory now draws fifty entries and offers more, which carries on from the last name, since a tree is in name order and stays in it.

The file view measured a six line file. Chapter 24 wants blame on every line, always, and chapter 25 gives the page 40kb, so the two together bound the page at about two hundred lines of source. Measured on an eight hundred line file it is 155kb. The page now fits what it can afford, counting each line rather than capping a count, because one long line costs more than one short one. Below the last line it says 200 of 800 lines and links to the whole file.

Both numbers are visible product decisions that the book's own budget forces. Flagged rather than hidden.

Three pages measured for the first time, and three are over

The budget table names seven paths. Every page not on it has never been measured, which is how the log page reached 219kb and the run page reached 190kb. Five more were pointed at the big fixture, with a build that says a great deal recorded on its tip.

page time budget payload
runs 1ms 10ms 1kb
one run 27ms 10ms 13kb
the profile 32ms 10ms 1kb
one commit 23ms 20ms

The payloads are fine and the times are not, which is a different disease from the four pages before it. Those sent too much. These do too much.

Two costs came off the profile already. git symbolic-ref is a file read now, and git count-objects -v is a walk of the object directory: both were a process each, per repository listed, and a profile lists as many as the account owns. The language guess is cached under the tip it was read from, which cannot change under that name, so a thousand-file ls-tree happens once. 44ms to 32ms.

What is left is structural. A profile row costs a config load, a ref read, an object walk and a scan of every thread to count open proposals. The scan is cached per thread, so a repository with fifty threads is fifty small disk reads before the row can say 3 proposals. Chapter 25 says a page gets one or two git invocations; this page gets a handful per repository, and the mockup draws fourteen.

The honest fix is a per-repository summary cached and invalidated on push, so the profile reads one record per repository. That is real work and it is not this pass.

The rows are not in the table yet, deliberately. A row that fails on purpose turns a green suite red forever and stops it telling anybody anything. The measurements are here instead, so the number is written down and the work is visible rather than forgotten.

Signed commits, and what chapter 23A used to say

23A said signatures are never required, and now three repositories require them. The old sentence was "Never reject a push for being unsigned; that is a policy for the project, not for barerepo to enforce." The second half of that is still right, so the rule is a per-repository key rather than a server setting: [access] require_signed_commits, off unless a repository asks. What changed is that the project can now hold barerepo to its own policy instead of asking people to remember. barerepo/server, barerepo/cli and barerepo/runner set it, because those three distribute the program itself. Nothing else on the server is affected.

It checks the maths, not just the header. The first version only looked for a gpgsig header, which stops somebody forgetting and stops nobody who is trying. It verifies now, against an allowed_signers file written from the ssh keys accounts already published for authentication. Each entry carries namespaces="git", so a sign in signature cannot be replayed as a commit signature. The principal is a wildcard, because a commit names an address barerepo never issued and has no way to tie to an account.

Retiring a key does not delete it. DeleteKey used to remove the row, which would have invalidated every commit that key had ever signed the moment somebody rotated. It sets retired_at now. Live keys go to authorized_keys, every key ever published goes to allowed_signers, and a retired one carries valid-before. git checks a signature against the commit's own timestamp, so old work still verifies and new work signed by the retired key does not.

Two things about that timestamp cost an hour each. valid-before is exclusive, so it is written one second after the moment of retirement, or a commit made in the same second as the rotation is refused. And it is written in local time with no suffix: ssh-keygen reads a bare timestamp as local, and the Z form this OpenSSH build was given did not parse at all, which silently turned every constraint into a refusal. Both are covered by tests that fail if either is undone.

The range is what the ref gains, not what the repository gains. The first version walked rev-list <new> --not --all, which skips any commit already in the repository. An unsigned commit pushed to refs/proposals/N is already in the repository, so landing it on master would have passed unread. It walks <new> --not <old> now, and a whole history on a branch that did not exist before.

source · rawbarerepo 0.1.0