feat: pair and similarity head with 0 layers have no layers - #383
Conversation
A pair similarity model with n_layers=0 whose output dimension equals the embedding dimension now has an empty head, so the embeddings are trained and used as is, and to_static_model gives the same embeddings as the trained model. Classifiers and regressors always keep their linear layer. MLPHead and the ONNX pipeline wrapper now pass the input through unchanged when the head has no layers, instead of failing to unpack an empty list.
Codecov Report✅ All modified and coverable lines are covered by tests.
🚀 New features to boost your workflow:
|
|
| if self.n_layers == 0 and self.embed_dim == self.out_dim: | ||
| return nn.Sequential() |
There was a problem hiding this comment.
Headless training loses gradient path When
n_layers=0, the dimensions match, and both freeze=True and freeze_weights=True, this empty head leaves nothing trainable. The training loop still calls loss.backward(), but the loss has no gradient path, so fitting fails. Previously, the linear head allowed this head-only training configuration to work.
There was a problem hiding this comment.
Hmm user error? They would need to set all of these options explicitly.
There was a problem hiding this comment.
That’s fair for the new headless configuration: explicitly setting both freeze=True and freeze_weights=True means the caller has asked for every parameter to be frozen, so there is no meaningful optimization to perform. I’d still avoid letting this fail later at loss.backward() with an opaque “does not require grad” error, because this combination worked before when the linear head remained trainable. A small validation in the training path (or at initialization) that raises a clear error when the model has no trainable parameters would make the behavior intentional and actionable; otherwise this is reasonable to treat as a caller error rather than a problem with the empty head.
Tip: You can customize Greptile's behavior for this repo with .greptile/rules.md and .greptile/config.json.
This PR adds an option to make models whose dimensionality matches their output targets have no head at all. This improves training, since we remove the linear layer after training anyway.
Here's the conditions:
for similarity, regression, and pairwise trainers, if you pass
n_layers==0AND the target and embedding dim are equal, we use no layer at all.