Skip to content

YOLO26 nano semantic segmentation

yolo26n_sem gives each image pixel a class score. For example, a two-class model can separate people from the background. It returns full image masks, not separate masks for individual people.

This implementation follows the nano backbone, top-down feature fusion, and semantic classifiers specified by YOLO26. The architecture reference is pinned to upstream revision abd16e0. The PyTorch implementation was written independently. Upstream checkpoints cannot be loaded directly. Holocron does not provide pretrained weights.

import torch
from holocron.models.segmentation import yolo26n_sem

model = yolo26n_sem(num_classes=19).eval()
with torch.inference_mode():
    scores = model(torch.rand(1, 3, 128, 192))
    labels = scores.argmax(dim=1)  # [1, 128, 192]

The model accepts rectangular images and odd image dimensions. It resizes scores to the exact input height and width. In training mode, the default model returns {"out": main_scores, "aux": auxiliary_scores}. Both have shape [batch, classes, height, width]. In evaluation mode, it returns only the main score tensor. Set auxiliary=False to return a tensor in both modes.

flowchart LR
    A[Image] --> B[Nano backbone]
    B --> C[Top-down feature fusion]
    C --> D[Main classifier at stride 8]
    C --> E[Auxiliary classifier at stride 16]
    D --> F[Resize to input dimensions]
    E --> G[Training loss only]
    F --> H[Class scores for each pixel]

Training

The existing segmentation trainer adds the main cross-entropy loss and 0.5 times the auxiliary loss. Labels equal to 255 are ignored. See the training guide for the VOC training command and reproducible learning checks.

The checked-in experiments use random initial weights. The synthetic run has separate training and validation images. The real-image run uses 120 Penn-Fudan images for training, 25 for model selection, and 25 for the final test. These checks show that the model learns. They do not reproduce Cityscapes or ADE20K accuracy, or establish hardware throughput.

Learning check Held-out result
Synthetic shapes, 24 validation images 95.02% mean IoU
Penn-Fudan, 25 test images 76.01% mean IoU, 59.94% person IoU

Deployment

For 19 classes, the model has 1,632,902 parameters during training. After folding convolution/BatchNorm pairs and removing the auxiliary head, it has 1,552,795 parameters, which rounds to the published 1.6 million. The larger parameter number in the upstream YAML summary comment is not the count of this semantic model.

model.eval().fuse()  # In-place conversion for inference
torch.save(model.state_dict(), "semantic-fused.pth")

loaded = yolo26n_sem(num_classes=19).eval().fuse()
loaded.load_state_dict(torch.load("semantic-fused.pth", weights_only=True))

Construct and fuse the destination model before loading fused weights. Keep an unfused checkpoint if further training is needed. CPU tests check fused prediction parity, saved-state reload, ONNX export and ONNX reference evaluator parity. CUDA, TensorRT, quantization, and upstream weight conversion are not validated here.

YOLO26Semantic

YOLO26Semantic(num_classes: int = 19, in_channels: int = 3, auxiliary: bool = True)

YOLO26 nano model for pixel classification.

The backbone and top-down neck follow the published nano architecture. A stride-eight classifier predicts the main mask. An auxiliary classifier supervises stride-sixteen features during training. Both outputs are resized to the exact input dimensions for Holocron's segmentation trainer.

This is an independent implementation. Upstream checkpoints are not supported, and no Cityscapes accuracy is claimed for these random weights.

PARAMETER DESCRIPTION
num_classes

number of output classes

TYPE: int DEFAULT: 19

in_channels

number of input image channels

TYPE: int DEFAULT: 3

auxiliary

include the training-only auxiliary classifier

TYPE: bool DEFAULT: True

RAISES DESCRIPTION
ValueError

if the number of classes or input channels is not positive

Source code in holocron/models/segmentation/yolo26.py
def __init__(self, num_classes: int = 19, in_channels: int = 3, auxiliary: bool = True) -> None:
    super().__init__()
    if num_classes < 1 or in_channels < 1:
        raise ValueError("num_classes and in_channels must be positive")
    self.num_classes = num_classes
    self.backbone = YOLO26Backbone(in_channels=in_channels)
    self.neck_p4 = C3k2(384, 128, use_c3k=True)
    self.neck_p3 = C3k2(256, 64, use_c3k=True)
    self.classifier = nn.Sequential(ConvNormAct(64, 64), nn.Conv2d(64, num_classes, 1))
    self.aux_classifier = nn.Sequential(ConvNormAct(128, 64), nn.Conv2d(64, num_classes, 1)) if auxiliary else None

yolo26n_sem

yolo26n_sem(pretrained: bool = False, progress: bool = True, **kwargs: Any) -> YOLO26Semantic

Build the YOLO26 nano semantic segmentation model.

PARAMETER DESCRIPTION
pretrained

request pretrained weights, which are not available

TYPE: bool DEFAULT: False

progress

kept for compatibility with the model factory API

TYPE: bool DEFAULT: True

**kwargs

arguments of :class:YOLO26Semantic

TYPE: Any DEFAULT: {}

RETURNS DESCRIPTION
YOLO26Semantic

an untrained semantic segmentation model

RAISES DESCRIPTION
ValueError

if pretrained weights are requested

Source code in holocron/models/segmentation/yolo26.py
def yolo26n_sem(pretrained: bool = False, progress: bool = True, **kwargs: Any) -> YOLO26Semantic:
    """Build the YOLO26 nano semantic segmentation model.

    Args:
        pretrained: request pretrained weights, which are not available
        progress: kept for compatibility with the model factory API
        **kwargs: arguments of :class:`YOLO26Semantic`

    Returns:
        an untrained semantic segmentation model

    Raises:
        ValueError: if pretrained weights are requested
    """
    if pretrained:
        raise ValueError("Pretrained YOLO26 semantic weights are not available in Holocron")
    return YOLO26Semantic(**kwargs)