Notes
Notes
We write rarely, and only about work we have done with our own hands. These pages are the studio thinking in public, dated in 2026, with no newsletter attached.
14 February 2026
Why we start with a hundred real questions
A demonstration prompt is a kind of theatre. Someone types a question the tool was prepared for, the answer arrives in a tidy paragraph, and the room relaxes. The week after the demonstration, a clerk asks the hundred questions that actually appear in the job, including the ones that are badly spelled, half-remembered, or aimed at a file that was never scanned. Those hundred are the brief.
We ask for them before we write much code. They come from mail, from tickets, from the last month of “have you got a minute.” We sit with the person who would have to live with a wrong answer and we write down what a right one would contain, including the cases where the only honest reply is that the file does not say. The set is ugly. It is also the only way we know whether retrieval is even possible.
A hundred is a lower bound. Some builds want three hundred because the types of request keep branching. We stop adding when a new question stops teaching us anything about a missing folder or a missing field. The extra questions after that are comfort, and comfort is how evaluation sets rot.
If an organisation cannot produce the questions, we treat that as a finding. It usually means the work is still in people’s heads. In that case the first job is writing things down, and a model will not do that writing for you.
3 March 2026
Retrieval beats a bigger model more often than you’d think
When a team asks for a more capable model, they are often asking for a missing page. The current model sounds vague because it has not been shown the protocol, or it has been shown a scan that a search step skipped. Swapping the model can hide that for a week. It does not fix the skip.
On the builds we take, the archive is the asset. A smaller model that is handed the right two pages will usually answer more calmly than a larger model that is handed a guess. The work, then, is the dull work: how files are named, which duplicates to ignore, how to keep a person’s permissions around the search, how to admit that a clause is not in the set.
We still choose a model. We just choose it after the retrieval path is honest. The evaluation set makes the conversation short. If a larger model lifts the score in a way the people in the room care about, we take it. If it only makes the prose prettier while still missing the page, we leave it.
This is also why we are slow to “just add the whole drive.” A drive is not a corpus. It is a mix of drafts, personal notes and the one signed PDF that actually governs the work. Retrieval that cannot tell those apart will sound informed and be wrong.
29 March 2026
What a document pipeline costs you in attention
People talk about pipelines as if the gain were the messages that file themselves. The real design is about the messages that should not. Every automated row you accept is a row nobody looked at. That can be fine for a booking date that matches a known pattern. It is not fine for a cancellation that names the wrong shipment.
Attention is the scarce thing on a dispatch desk or in a clinic office. A pipeline that dumps every low-confidence case into a heap has only moved the heap. A pipeline that hides low confidence has spent the attention you will need later, when a customer rings. We spend the discovery week watching which exceptions a person actually wants to see, and we make that queue short and visible.
There is a cost in the other direction too. If you insist that a person tick every field, you have bought a typing assistant and called it automation. The honest conversation is which fields, on which types, are allowed to pass, and who is named when one of those passes is wrong.
We write that down. It looks like a dull table. Six months later it is the only artefact anyone can use to argue with a new manager who wants the queue switched off.
18 April 2026
Choosing between a hosted API and a model on your own box
The question is usually asked as a matter of principle. In the work we do it is a matter of where the files are allowed to go. If a contract cannot leave Singapore, or cannot leave the building, then a hosted interface in another region is the wrong tool no matter how convenient the demo. If the files are already in a vendor’s region under your existing rules, a hosted interface can be the simpler path.
A box next to the archive has costs that a slide rarely mentions: someone has to keep it updated, the room has to stay cool, and a model that fitted last quarter may not fit the next one. It also has a virtue the slide does mention, and that virtue is real: the bytes do not take a walk. We will argue for the box when the documents are the sort that would be painful to explain on a breach form.
A hosted interface has a different quiet cost. Prices move. Models are renamed. A prompt that was fine in March is charged differently in June. That is why usage is billed at cost, with the records attached, and why the evaluation set is re-run when a vendor announces a change. You can see the number. You can cap it. You can switch.
We do not have a house style. Discovery ends with a recommendation that follows the files, the rules you already have, and the person who will have to live with the machine when we have gone back to Kampong Glam.
6 May 2026
The quiet part: what breaks three months later
Hand-over day is a good day. The scores are green, the guide is printed, someone on the team can add a file. Three months later the thing that breaks is rarely the model. It is the world around the model. A folder is reorganised. A template gains a field. The person who owned the evaluation set takes another job. A vendor ships a “silent” update that changes a refusal into a guess.
The failures look polite. An assistant still answers. A pipeline still writes rows. The rows are slightly wrong in a way that only the person who left would have spotted. This is why we argue for a named owner on your side, and for a set of questions that can be re-run without us. A monthly visit is one way to keep that habit. A calendar reminder and a script is another. What does not work is hope.
The other quiet break is permission. Someone shares a drive more widely “so the assistant can see it,” and a month later the assistant can see a draft it was never meant to cite. Access should follow the folders you already trust. When that rule is bent, we would rather hear about it than discover it in a log.
None of this is dramatic. It is housekeeping. The studio exists for the builds; the housekeeping is how a build stays a tool instead of becoming a story people tell about the month it worked.
Desk
Have a question about one of these? Write to us.
If a note maps onto a pile of files you already have, send a few sentences. We will tell you whether it looks like a discovery week. The photograph is the corner of the desk where those sentences get read.
