Text recognition¶
Holocron provides an experimental shared character backbone, glyph classifier,
and CNN/BiGRU recognizer trained with connectionist temporal classification
(CTC). Import them from holocron.models.recognition; training and synthetic
benchmark generation are provided by the
recognition reference.
The models accept grayscale tensors of shape (N, 1, 32, W), normalized to
[-1, 1], with white padding equal to +1 and width divisible by four.
lengths contains each image's unpadded width divided by four. The recognizer
returns log probabilities of shape (W // 4, N, alphabet_size + 1) with CTC
blank at index zero. Loss and decoding must use the true lengths.
import torch
from holocron.models.recognition import CharacterClassifier, CTCRecognizer
from holocron.utils import CTCCodec
codec = CTCCodec("AB12 .")
classifier = CharacterClassifier(num_classes=len(codec.alphabet))
reader = CTCRecognizer(num_classes=len(codec.alphabet))
# After glyph pretraining, transfer the visual encoder into the line model.
reader.backbone.load_state_dict(classifier.backbone.state_dict())
# Example input contract; train or load weights before using predictions.
images = torch.ones(2, 1, 32, 80)
lengths = torch.tensor([20, 12])
reader.eval()
with torch.inference_mode():
log_probabilities = reader(images, lengths)
labels = log_probabilities.argmax(-1).transpose(0, 1)
transcripts = [codec.decode(row[:length].tolist()) for row, length in zip(labels, lengths)]
CTCCodec accepts a custom printable Unicode alphabet, including ordinary
space. It preserves predicted punctuation and repeats separated by blanks.
holocron.utils.prefix_beam_decode optionally sums alternative alignments;
it uses no dictionary or language model. Greedy decoding remains the measured
default.
The synthetic curriculum benchmarks 83 symbols across ten font families, holding two families out of every training stage. Its final degraded-page CER is 1.06% on both seen and held-out families, with 64% and 53% exact pages. The full report includes the control, training budgets, seeds, and residual errors. No public pretrained checkpoint is downloaded. These measurements cover separated, horizontal, single-column synthetic pages; real scans and general layouts remain unvalidated. Page detection and synthetic rendering remain experimental reference code.
CharacterBackbone
¶
Shared grayscale encoder for glyph classification and line recognition.
Input images have shape (N, 1, 32, W) with widths divisible by four.
The output has shape (N, 96, 4, W // 4). Normalize grayscale images
to [-1, 1], using +1 for white padding.
Source code in holocron/models/recognition/crnn.py
CharacterClassifier
¶
CharacterClassifier(num_classes: int)
Classify glyphs while retaining their vertical baseline position.
| PARAMETER | DESCRIPTION |
|---|---|
num_classes
|
number of output character classes
TYPE:
|
The input contract matches :class:CharacterBackbone. lengths is a
one-dimensional integer tensor containing each image's unpadded width
divided by four. Returns logits of shape (N, num_classes).
Transfer backbone.state_dict() directly to :class:CTCRecognizer.
Source code in holocron/models/recognition/crnn.py
CTCRecognizer
¶
CTCRecognizer(num_classes: int)
Compact CNN/BiGRU recognizer for variable-width text lines.
| PARAMETER | DESCRIPTION |
|---|---|
num_classes
|
alphabet size excluding the CTC blank
TYPE:
|
The input contract matches :class:CharacterBackbone. lengths is a
one-dimensional integer tensor of unpadded widths divided by four.
Packed recurrence excludes padding from bidirectional context. Returns
log probabilities of shape (W // 4, N, num_classes + 1) with blank at
index zero. Ignore frames beyond each length during loss and decoding.
The architecture is experimental; no pretrained weights are downloaded.