How VoiceMark reads, and what it refuses to do.
The short version is on the VoiceMark page. This is the long one: the problem, the method, the complete refusals, the cases where it reads badly, the research it was built against, and where the validation currently stands.
The problem
The failure is not that the writing is wrong.
Prose that comes out of a model is usually accurate enough. What it does is perform the appearance of thought on the page, in place of recording thought that has already happened. That difference leaves a shape in the writing, and the shape can be read.
Reading it is not the same as proving it. A pattern in a text tells you something about the text. It does not tell you who sat at the keyboard, and it never tells you why. VoiceMark is built around that limit rather than around a claim that cannot be supported.
Three things, and it says what each one is worth.
It reads the prose, not the keystrokes.
Typing rhythm, paste history and draft replay record that a person operated a keyboard. Work published in January 2026 showed formally that they carry no information about who originated the words, and that they can be reproduced at will by anyone who wants to. VoiceMark reads properties of the writing itself, so a finding is a fact about the text in front of you.
Read the paperTwo readings, held apart.
One says how strong the evidence is. The other says what that evidence points to. Keeping them separate stops a weak signal being reported as a firm conclusion, which is where most of the harm in this field has come from.
Every finding says what it leaves open.
A reading names the passage, says what it found there, and says what would change the reading. Where the evidence runs out, the report says so rather than filling the gap.
What this will not tell you.
Everyone in this category publishes what their tool can do. This is the other half, and it is on the page rather than in a footnote.
-
It will not give you a percentage.
A number implies a precision this method does not have, and the ombudsman for higher education in England and Wales, the exam boards' joint council and the sector's own technology body have each said that a tool's output cannot carry a finding on its own.
-
It will not tell you that a piece of writing is human.
A clean reading means the patterns we look for are not present in this text. Absence is not proof, and we will not report it as though it were.
-
It will not be your evidence.
Not in a disciplinary hearing, an academic integrity case, an employment matter or a publishing dispute. A reading is one input to a decision a named person takes, and it is not a substitute for that person doing the work.
-
It will not help you get past a detector.
Cleaning a text of the patterns we name is evidence that the text was revised. It is not evidence that a person wrote it, and we do not sell it as such.
The cases where a confident tool does the most damage.
These sit here rather than in a footnote because they are the cases where a confident tool does the most damage, and because you are entitled to know them before you rely on anything we say.
Four things arrived at the same answer, and none of them were ours.
We did not set out to build a cautious product. The caution arrived from four directions at once.
-
January 2026
Process evidence does not answer the question.
A formal result showed that keystroke and timing evidence carries no information about who originated a text, and that it can be forged at will.
arXiv 2601.17280 Preprint, not yet peer reviewed. -
March 2026
The uneven error is structural.
An argument that writers whose prose resembles machine output face a higher rate of false flags whatever the tool does, so a better model does not remove it.
arXiv 2603.20254 Preprint, not yet peer reviewed. -
August 2026
Most flags from a percentage tool are wrong.
A demonstration that at the rates these tools actually run at, the majority of what they flag is not what they say it is.
Read the paper Preprint, not yet peer reviewed. -
Across the same period
The bodies that decide these cases said the same thing.
The ombudsman for higher education, the exam boards' joint council and the sector's own technology body each concluded that a tool's output cannot be the basis of a finding on its own.
The validation position, stated plainly.
The research above is the reasoning the method was built against, not a validation of VoiceMark itself. Three of those four sources are preprints and are labelled as such. No independent evaluation of VoiceMark has been published, and until one has been we will not describe the method as validated.
What we can say is narrower and it is what the refusals above describe: a reading reports marker-grounded confidence with the evidence attached, it does not produce a score or a verdict, and it is not sufficient on its own for a decision about a person. If that changes, this section changes with it and will say who did the work.
That is the method. The product is on the other page.
Pricing, the mirror you can try without an account, and what happens to what you write.