TIMELESS TALES CLASSICSTimeless tales. New voices.

← All entries

Entry 003 · 15 August 2026 · Method · Audiobooks

Ten out of ten estimates were too high

Estimate too high
10 of 10Character/volume combinations across volumes 1 to 4, against the hand-checked result
Largest error factor
21×Osko in volume 4: 1.9% estimated, 0.09% actual
Machine labels
205 to 38OSKO blocks against mentions of Osko in the source text; 27 remained
Heard wrongly by
nobodyFound before narration — no volume was published this way

Before a Karl May volume is narrated, somebody has to decide which character gets a voice of their own. In Through the Desert it is three: the narrator, Halef and Lord Lindsay. Everyone else is spoken by the narrator.

That decision rests on each character’s share of the dialogue, and nobody counts that by hand through fifteen hours of text. A script does it: it breaks the novel into blocks and writes above each one who is speaking. The blocks give the share, the share gives the voice.

The script is wrong. And not at random.

The measurement

We held the machine estimate for volumes 1 to 4 against what the hand-check actually produced — ten character/volume combinations.

In ten cases out of ten, the estimate was above the true share. Not above sometimes and below at others: always above, by a factor of 1.6 to 21.

The worst case

Osko, a companion in volumes 4 to 6. In volume 4 the script assigned him 205 blocks. The source text names him 38 times. After the hand-check, 27 blocks remained in which he actually speaks.

An estimated 1.9 per cent share became a measured 0.09 per cent — a factor of twenty-one.

Why the machine always errs the same way

Because it has to guess. Dialogue in an 1892 novel is rarely tagged cleanly; often there is only “he said”, and who he is follows from context. The heuristic assigns such blocks to the nearest plausible name — and the nearest plausible name is that of the character currently in view.

A character label thereby becomes, in effect, a placeholder for “speaker unclear”. The error does run both ways — real dialogue sometimes stays buried, unmarked, inside narrative blocks — but the overestimate dominated in every single volume.

The rule that follows

It is one-directional, and that is its entire value:

What the segmenter sees below the threshold is certainly below it — no hand-check needed there. A high estimate, by contrast, proves nothing.

That saves the work where it yields nothing and leaves it standing where it decides. A practical cross-check costs two minutes: hold the number of name mentions in the source against the number of labels assigned. 205 against 38 says everything at once.

The counter-case that nearly produced a second mistake

Once we had learned to distrust the segmenter, the inverse seemed obvious: a chapter with one single narrator block and no Halef lines must surely be a segmentation that was forgotten.

It is not. Volume 4 has eight such parts — and in all of them Kara Ben Nemsi is travelling alone. The text says so itself: “waited in vain for Halef”, “I do not see my companions”. A single narrator block is correct there.

Had we given Halef a voice because a warning was red, we would have had a character speak who is not in the room. That would no longer be a process error but unfaithful to the work — and noticed by nobody except a reader who knows the book.

So before forcing a character voice, always check the character’s presence in the source.

What never happened

Nobody heard this error. Everything described was found before narration, while proof-reading the scripts. No volume was published with voices wrongly assigned.

That is not luck but the order of work: read against the text first, render second. Every regeneration costs the same credits again — a mistake that slips through into the recording is not embarrassing, it is expensive.

Never miss one

New entries by email

One list for new logbook entries and new releases — we send nothing else. If you only want the entries, take the RSS feed.