What we publish
We publish method, not papers
We have not published academic papers, and we are not going to imply otherwise. What we do publish is the working material behind our projects: how we build evaluation sets, what we measured, and where a method stopped paying. These notes are being written now.
Nothing here is gated, and nothing here claims academic peer review we have not done.

A method is only useful here if it survives contact with a client’s real, incomplete data.
What is being prepared
Four summaries are in progress. Each states the method, the dataset it was validated on, and the honest limit of where it transfers.
How we build an evaluation set before we build a model
The small, noisy, partly mislabelled datasets a mid-sized business actually has — and how we turn them into a test set that can fail a model before a client pays for one.
Evaluating a model where a confident wrong answer is expensive
Why accuracy is the wrong headline number when the cost of the two error types differs, and what we report instead on every project.
Evaluation sets for Bulgarian-language document extraction
How we assemble a test set from real invoices, contracts and delivery notes, including the near-duplicates, bad scans and mixed-script pages a public benchmark quietly excludes.
Open-weight models under quantization
Measured accuracy loss on structured extraction as precision drops, and the point at which the cheaper deployment stops being cheaper once review time is counted.
Working on something adjacent, or want a paper before the summary is written? Write to us.

Have a problem that looks like a research question?
Some of them are. If a problem needs a study rather than a build, we will say so, scope the study, and tell you what a negative result would look like.
We publish under our own names, and we cite the work we build on.