Skip to content

Methodology

TorchScan describes one executed model call or workload. It combines two observation mechanisms while keeping their results separate:

  1. Module hooks capture hierarchy, call order, tensor metadata, parameters, buffers, and module formulas.
  2. PyTorch operator dispatch captures executed operations with registered FLOP formulas.

Neither mechanism is a hardware benchmark or a complete graph export.

Module analysis

crawl_module and summary make one forward call. Generated inputs add a batch dimension of one; caller-provided args and kwargs are forwarded unchanged. Analysis runs in evaluation mode with gradients disabled, then restores each module's previous training flag.

Use args and kwargs when the model does not accept a leading batch dimension, including batch_first=False sequence modules. Calls sharing the same module instance must be serialized because crawling temporarily changes its training state and installs hooks.

Layer identity is the full module path plus a call index. This distinguishes repeated calls through a shared module without inventing duplicate parameters.

Supported module formulas calculate theoretical FLOPs, MACs, DMAs, and receptive field. Unsupported work changes the affected metric state instead of contributing zero. Formula definitions and their tested boundaries live in the package source; diagnostics expose unsupported paths.

Operator FLOPs

measure_flops uses PyTorch's torch.utils.flop_counter.FlopCounterMode around one owner-provided workload. The dispatcher can observe functional operations and work inside custom modules that hooks cannot assign to a supported leaf formula.

For structure, shapes, parameters, and buffers alone, use mode="structure" in crawl_module or summary. This skips operator counting and module formulas while preserving one evaluation forward pass. Skipped compute totals are explicitly unavailable with method not_requested; strict checks apply only to requested metrics. Module formula work in full mode runs in post-hooks with dispatch suspended, allowing activations to be released during execution instead of retaining them for deferred analysis.

Only operators with registered or caller-provided formulas contribute to the known count. TorchScan records executed but uncounted operators and marks the result partial. Caller formulas are scoped to one invocation and use the installed PyTorch version's shape-formula contract.

The workload owns model state, gradient mode, autocast, device placement, warmup, and side effects. Exceptions are propagated unchanged.

Why the two FLOP views can differ

Module formulas and operator formulas may choose different boundaries or arithmetic conventions. Composite modules can decompose into several dispatcher operations, while fused operators can combine work that a module formula describes separately. Custom formulas can also use experiment-specific conventions.

For this reason TorchScan labels both methods and does not reconcile them into one number. A paper or regression report should state which view it uses.

Partial is a lower bound, not coverage

For a partial result, known_value is the sum of counted work. It is a lower bound only. TorchScan does not calculate a coverage percentage because operator kinds or call counts do not reveal the cost of the missing work.

Diagnostics are part of the measurement. Store them with the numeric result and resolve them before making a completeness claim.

Peak memory

measure_peak_memory invokes a zero-argument workload once. CPU uses profiler memory categories; supported accelerators use allocator statistics. These are PyTorch measurements, and concurrent allocations can affect them. See Peak memory for backend boundaries and comparison requirements.

Latency

Use torch.utils.benchmark.Timer for warmup, replicates, and synchronization. Keep latency separate from theoretical operation counts.

Minimum reproducibility record

Follow the reproducible reporting checklist: retain the report, diagnostics, software versions, model revision, input metadata, and custom formulas. Workload measurements also need hardware and execution-state details. Target-device acceptance requires real checks on that device.

FLOP conventions

torchscan_flops_v1 keeps module/operator counts separate. Priority: caller override > native PyTorch > invocation-only fallback; no global mutation. N = elements, R = rows, C = channels, A = affine tensors present (0–2). Scalar arithmetic, exp/sqrt, comparison, and selection each cost one operation.

Work Count / boundary
K-term dot, K>0 Module: K multiplies + K-1 adds; native: 2K. Empty dots: zero.
Bias/scaling Module: one bias add/output. Native fused bias and alpha/beta scaling omitted; separate arithmetic counts.
Grouped convolution Outputs × Cin/groups × kernel_volume terms; dense padded work.
Transposed convolution 2 × input_elements × Cout/groups × kernel_volume, plus module bias; cropped scatter products included.
Fixed pooling, K values/window Max: K-1; average: K. Nominal dense padding/ceil windows.
Broadcast/reduction Per output; sum: max(N-R,0); mean adds R divides.
Stable/safe softmax 5N-2R: max, subtract, exp, sum, divide. Safe adds 2N comparisons/selections.
LayerNorm/GroupNorm 6N+2R+AN: mean N, variance 3N, normalize 2N, eps/sqrt 2R, affine AN.
Rows LayerNorm: N/prod(normalized_shape); GroupNorm: batch × groups. Empty rows/groups unsupported; zero batches supported.
BatchNorm Saved stats: 2N+2C+AN; batch stats add 4N. Actual buffers select the path; passed mean/variance updates add 3C/5C. Empty modules: zero; empty native calls unsupported.
Dropout Eval/p=0: zero; training mask/rescale: 2N, or N at p=1. RNG excluded.

Transpose example: (2,4,3), Cout=6, groups=2, kernel=3, stride=2, padding=output_padding=1: 24×3×3×2 + 72 biases = 504 module FLOPs. Shape options add no scatter MACs.

Attention Included work / limits
Batched MHA Both layouts, unequal widths: projections/bias, Q scale, dense products, softmax, masks, dropout, optional head averaging.
CPU SDPA fallback Dense 4D, matching batch/heads and Q/K widths, equal K/V lengths, no dropout: native matrix work + score scale + softmax + masks.
Masks/math Each explicit/causal mask costs one operation/score; boolean conversion counts stored entries. Dense products unchanged. Math can scale Q/K separately and use safe softmax; explicit scale changes values only.
Native fused attention Matrix core only; ancillary gaps keep counts partial, including PyTorch 2.1 CPU.

Reduction dtype: dtype > out.dtype > input. Fallback integer/boolean arithmetic, views/copies/fills/allocation, and Python constants are excluded. Native matrix counts include integers. Complex module/fallback arithmetic and sparse/nested work stay incomplete; native complex counts remain partial. Mixed-call diagnostics and strict mode are preserved.

Remaining gaps: MHA unbatched/empty sequences, add_bias_kv/add_zero_attn, fused MHA/encoder, specialized attention, unknown normalization/softmax backward, RNG/optimizer/embedding/gather, and unregistered activations/pooling. Transformer requires native stacks, ReLU, final LayerNorm/Identity/None. Adaptive/other pooling metrics retain legacy approximations. Normalization kernel algorithms can differ; CPU/meta checks do not validate CUDA/MPS or latency.

crawl_module: eval/no_grad forward. measure_flops: supplied forward/backward work, without a backward multiplier. Run python scripts/benchmark.py --json /tmp/torchscan-matrix.json: CPU, seed=0, one thread, float32 (1,3,32,32). Cells retain complete/partial (>=)/unavailable states and JSON diagnostics. Derivations: tests/test_flops.py; seven no-weight-download integration smokes: tests/test_model_zoo.py.