Testing an IF story is Hard
I'm on the verge of updating Chord to 4.0.0 with changes inspired or more accurately exposed by trying to port an old game that no one has ever played outside of my high school.
I'd talked about the problem in my last post. As I was making those changes, I also became ensorcelled by the possibility of a truly cross-platform IDE built on Avalonia. In the macOS IDE and this new IDE the one thing that I am still very frustrated with is testing and I've spent way too much time fussing over it. It's a tease, just out of reach of something really cool.
The current IDE testing implementation is based on the concept of a branching tree, which is obviously where Graham landed with Inform 7, but I've been chasing something crazier.

Sharpee's world model is data, so a running story can be introspected regardless of the language that wrote it. Chord is self-contained, so a Chord story can also be introspected completely and statically: nothing a story does is hidden from the compiler or the IDE.
This was not true prior to v4 since I had allowed the language to "hatch" to Typescript includes. I decided to remove the hatch mechanism and leave those use cases to either feature requests or Chord extensions.
This allowed the IDE to know where everything is or The Whole Truth. this includes:
- the Index tab lists every room, region, thing, person, action, and phrase, and every row opens at the line that declared it
- the build report prints the story's numbers, counted from the compiled story rather than guessed from the source
- every problem names its file and line, and there is no longer a problem of the form "there is code here I can't see into"
- the Testing tab asserts on channel output, and since nothing a story does lives outside the story, those assertions can reach everything it says
Back to the rabbit (rat) hole. I ran the port of Secret Letter in the IDE and it takes about 80 seconds to load the tree of tests. Obviously that's completely unacceptable. So I started pushing Claude to find ways to test from the moment the author opens the IDE and starts a new story.
This led to the creation of compile time "lenses" that include:
- the world index: a map, what the player can actually reach (a locked door counts as passable only once the thing that opens it is itself reachable), and what the prose names that the story never defines
- the derived rule suite: every rule I write is a precondition, an action, and effects, so the compiler can enumerate one test per clause branch. The list is complete, so coverage has a real denominator. The tests themselves run through the real engine, and a failing one fails the build
- declared states nothing assigns: values no rule ever writes and dimensions no rule ever reads, read straight off the compiled story without running anything
There were limits to these kinds of compile-time lenses:
- the compiled story records what I wrote, not what the platform does with it. A door connects two rooms and starts locked, but I never wrote "locked". The first state I list is the initial state, but no statement assigns it. An exit north gets its return exit south for free. A tool that reads only my words reports all three as bugs, so every lens has to carry the platform's defaults beside the record
- a story models possibilities, not actualities. The compiled story says what can happen, never in what order, and whether an ending is reachable is a question about a path through states. I insisted on the combinatorial test anyway. On a small story it ran its 15-minute budget, saw 23,163 states through 715,903 commands, and never found a 29-command winning path that is written down in the repo
- the parser and the prose both sit outside the compiled story. "The tent flap hangs open" names a thing in the player's head, and whether it is a thing in my story is guesswork over English. Even when the thing exists, whether the engine answers to "pole" or only to "post" is known only by asking the engine. So those checks have to type commands at the real engine, and they ship as candidates and findings, never as errors.
This still frustrates me. I still feel like there should be a way to provide complete deterministic testing through to all possible endings. Again, I'm stupid.
We still end up somewhere in the middle of where I'd like to be. The compiled story cannot tell me an ending is reachable, so I supply that part myself: when I play through to an ending in the Testing tab, that line is the proof, and it replays byte for byte at the pinned seed on every run. The tool owns everything below the line. Every rule I wrote gets its own test, arranged directly into the state it needs rather than played to, and the coverage report names every branch that was never run. Complete and deterministic, with the one requirement that the author provides the endings. The lenses sit beside that as diagnostics
And in the course of writing this blog, I was inspired to discuss my wishes again with Claude and we may have found the syntax for deterministic testing that avoid combinatorial explosion performance.
We add authored sets of claims.
claim the story can be won
ending: victory
needs the Iron Gates, the Gravel Drive, the Fountain Court, the Boiler Shed, the Entrance Hall, the Kitchen, the Greenhouse, the Folly Hill, the Folly
needs Tobias, the stopcock, the primer plunger, the boiler, the garden shears, the sherry bottle, Mrs Kettle, the vine, the silver locket, the folly door, the fuse, the deed box, the deed
needs ask, turn, push, turn on, take, give, prune, open, cut
claim the diary page has been read
needs the Iron Gates, the Gravel Drive, the Fountain Court, the Entrance Hall, the Kitchen, the Study
needs the sherry bottle, Mrs Kettle, the travelling trunk, the diary page
needs take, give, open, read
claim the deed box never leaves the Folly
needs the Folly, the Folly Hill, the Greenhouse, the Fountain Court, the Boiler Shed, the Gravel Drive, the Iron Gates
needs Tobias, the stopcock, the primer plunger, the boiler, the folly door, the deed box
needs ask, turn, push, turn on, open, takeThat syntax has been tested and works for a small story. A new ADR is completed and ready for implementation.
This still does not settle the IDE. I think the claims addition might help resolve that too.