On torch 2.8, exporting a transformer with torch.onnx.export(..., dynamo=True) and named torch.export.Dim
objects prints a ConstraintViolationError and falls back to the slowest way of capturing the model. For a
ModernBERT decision model, that meant about four and a half minutes per export instead of about a minute. Naming
the dynamic dimensions with strings fixes it.
The error
torch.fx.experimental.symbolic_shapes.ConstraintViolationError: Constraints violated (batch, options, tokens)!
- Not all values of batch = L['input_ids'].size()[0] in the specified range satisfy the generated guard
L['input_ids'].size()[0] != 1.
...
- Not all values of tokens = L['input_ids'].size()[1] in the specified range satisfy the generated guard
1 != L['input_ids'].size()[1].
- Not all values of options = L['marker_positions'].size()[1] in the specified range satisfy the generated guard
1 != L['marker_positions'].size()[1].
It comes from the call the PyTorch docs show, with named dimensions:
batch, tokens, options = (torch.export.Dim(n) for n in ("batch", "tokens", "options"))
dynamic = {
"input_ids": {0: batch, 1: tokens},
"attention_mask": {0: batch, 1: tokens},
"marker_positions": {0: batch, 1: options},
"marker_mask": {0: batch, 1: options},
"question_type": {0: batch},
}
torch.onnx.export(model, args, dynamo=True, dynamic_shapes=dynamic, ...)
The export doesn't fail. torch.onnx.export(dynamo=True) tries its capture strategies in order, and only the last
one succeeds:
[torch.onnx] Obtain model graph for `Wrapper([...]` with `torch.export.export(..., strict=False)`... ❌
[torch.onnx] Obtain model graph for `Wrapper([...]` with `torch.export.export(..., strict=True)`... ❌
[torch.onnx] Obtain model graph for `Wrapper([...]` with `torch.export draft_export`... ✅
The cause
A named Dim declares a range, and on torch 2.8 that range includes 1. Transformer code often behaves differently
when a dimension is 1 (broadcasting masks, squeezing a dimension, special-casing a single row), so torch records a
guard such as batch != 1 that contradicts the declared range, and draft_export, the slow path, is the only
strategy that accepts it.
The fix: name the dimensions with strings
dynamic = {
"input_ids": {0: "batch", 1: "tokens"},
"attention_mask": {0: "batch", 1: "tokens"},
"marker_positions": {0: "batch", 1: "options"},
"marker_mask": {0: "batch", 1: "options"},
"question_type": {0: "batch"},
}
torch.onnx turns each string into torch.export.Dim.DYNAMIC, which declares no range and lets torch infer one,
and then names the ONNX graph's axes after your strings. The first strategy succeeds, with no constraint errors:
[torch.onnx] Obtain model graph for `Wrapper([...]` with `torch.export.export(..., strict=False)`... ✅
The export takes about a minute, and onnxruntime agreed with torch to 1.3×10⁻⁶ in probability. Keeping the names
matters if you run the graph on a provider that compiles for fixed shapes, such as Core ML: onnxruntime's
add_free_dimension_override_by_name("tokens", 128) finds a dimension by its name, so an export that names its
axes s0, s1 would break that.
Check the model, not just the export
An export can succeed, match the PyTorch model it was traced from, and still be wrong, because that PyTorch model can be wrong: transformers 4.57.6 computes ModernBERT differently from 5.x. So pin the library that defines the model (since opendxp 0.3.0, the converters refuse transformers older than 5.2), and compare the package with answers recorded by the model's own code:
pip install "opendxp[export,laya]" "onnxruntime-gpu==1.23.*" # converter, Laya's own code, onnxruntime for CUDA 12
opendxp export laya ./laya-checkpoint ./laya-package
opendxp conformance generate ./laya-package --native ./laya-checkpoint --runtime laya
opendxp check ./laya-package --device cuda
opendxp conformance generate records the answers the model's own code gives to 52 fixed requests. opendxp check replays them through the package and fails if any decision changes or any probability moves by more than
0.01. On a CPU, install "opendxp[export,laya,onnx]" instead and drop --device cuda. Record the conformance file
where you trust the model's own code, then check the package everywhere else. Our packages of Julia 1 (Supersonic
Labs) and three Laya checkpoints (Convai Innovations) pass 52 of 52 cases on an NVIDIA A40, and on an x86 server
CPU they match the answers recorded on a Mac to the fifth decimal.
Related: the OpenDXP docs, Publishing your first decision model.