In short
The Humanly platform suggests that we should judge who wrote a text—a human, AI, or both—not based on the final result, but on the recorded writing process. An analysis of why this is a paradigm shift and where it falls short.
The problem with detecting AI-generated text is that the final text reveals nothing about how it was created. A teacher, reviewer, or editor sees only the result and tries to guess—based on the style—whether it was written by hand, generated by AI, or a hybrid of the two. Humanly—a research platform from arXiv—proposes moving away from analyzing the product and toward certifying the process.
The idea is simple: the user writes in a controlled environment that records all activity and every interaction with the built-in AI. Upon completion, the session is packaged into a “sealed writing certificate” with anomaly checks that take the environment’s configuration into account. The certificate is not an assessment of the text’s quality, but rather evidence of exactly how it was produced.
This is a fundamentally different approach. Instead of statistical detectors that attempt to distinguish “AI text” from “human text” based on linguistic features—and consistently fail— Humanly frames the question differently: let’s record keystroke patterns, timings, and tool interventions, and then verify whether the process matches the stated scenario.
Red-teaming research shows that the built-in Typing Detector distinguishes between manual typing and automated typing simulation. But this is where it gets really interesting. The detector specifically catches automated input—macros, script-based insertions—rather than the actual use of AI as an intelligent assistant. If a person carefully retypes text generated by Claude character by character, the process appears “human,” and the certificate will confirm this. The platform verifies the method of input, not the source of the ideas.
For scenarios like term papers and peer review, this is still a useful development. The process certificate provides stronger evidence than any analyzer of the final text—but only within the scope of its threat model. It detects cheating at the mechanical level, not at the level of cognitive authorship.
This is a real trade-off that Humanly highlights but does not resolve: to prove the integrity of a piece of writing, you must turn the very act of writing into an observable process. This works for exams and certification, where oversight is already expected. But in personal or creative work, the cost is the loss of privacy regarding the thought process itself. The platform turns the process into proof, and that is precisely why it cannot be universal.
Source: cs.CL updates on arXiv.org