It would be useful if Leo was putting question marks in square brackets (or similar) for parts of the text that is difficult to transcribe or illegible. For now it seems to either skip chunk of text, or even worse, guess the words or entire lines. These guesses are quite confusing, because sometimes, they may seem “fitting” enough to appear in a certain document, but after reading the transcription and checking it, they appear to be completely made-up. This guessing may be dangerously misleading and is actually worse than having lines missed. Having a clear indication via symbol that Leo struggled with chucn of document would be better for later checking the transcription.
1 Like
Thanks for this! I see where you’re coming from with dangerously misleading transcriptions. In practice, Leo is guessing every single word of the transcript. For the model, the difference between legibility and illegibility is not binary but scalar. The reason why we don’t want to include an [illegible] sign is because we don’t want to limit the potential scope for the model to learn in the future. To address this issue we’re planning to add confidence metrics. See here: