Discovery week
Five days inside the way you already work. We sit with the people who touch the documents, read a slice of the archive, and watch where a morning actually goes. The week is long enough to see whether retrieval, a pipeline, or an evaluation set would help. We come back with a short paper: three options, a rough shape of effort, and a plain view on whether building anything is worth it. Sometimes the honest answer is to tidy the files and stop.
What you get at the end
A paper of a few pages, the three options, and a go or stop that we will stand behind on a call.
typically 1 week
Assistant over your own documents
A search and answer tool that stays on your contracts, protocols and internal rules. Each reply carries a pointer to the source so a person can open the page and check. Access follows the folders you already share. We keep a log of queries so you can see what is being asked, and so the evaluation set can grow from real use. The assistant does not draft new policy or speak for the organisation in public.
What you get at the end
A working assistant, source citations, access rules that match your folders, and a query log you can read.
4–8 weeks
Document and intake pipelines
Incoming PDFs, scans and mail often land as a heap that a person then types into a sheet. We extract the fields you already track, assign a type, and write a row into the table or CRM you use now. Doubtful cases — a scan that will not read, a message that matches no type — go to a short queue for a person. That queue is the point of the design, because exceptions are where a clerk’s week actually goes.
What you get at the end
A running intake path, a human queue, and a written map of the types and fields the path understands.
3–6 weeks
Evaluation harness
Before we change a prompt or a model, we want a number that means something to the people who will live with the answers. We gather one hundred to three hundred questions taken from real work, write an agreed answer with you, and score the model against that set. The same set is re-run when a file is added, a prompt is edited, or a vendor updates a model. Failures are readable: which question slipped, which source was missed.
What you get at the end
The question set, the scoring script, and a simple way to run it again after a change.
2–4 weeks
Hand-over and training
We sit for half a day with the people who will use the tool and the person who will keep the files in order. We walk through the ordinary path, the failure path, and the two or three knobs that are safe to turn. You leave with a written guide, access to the repository, and a list of the things that should wait for us. The session uses your documents, on your desks, with the questions that come up that morning.
What you get at the end
A live session, a written guide, repository access, and a short list of what to leave alone.
typically 1 week
Quiet retainer
Once a month we look at the same three things: whether a vendor model has moved, whether the evaluation scores have drifted, and whether token spend has jumped for a reason you would want to know. Small fixes sit inside that visit. Larger changes go back to a scoped piece of work. The retainer is optional. Some organisations run the harness themselves after hand-over. Others want a named day in the calendar. We keep the visit short and write down what we saw.
What you get at the end
A monthly note, the scores, a line on spend, and the small repairs that fit the visit.
typically 1 day a month