↓Skip to main content

Publications

* indicates equal contribution or alphabetic order. The arXiv version is preferred for the most up-to-date content.

Transferable Graph Metanetworks

Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar 

arXiv preprint.

We study size generalization in weight-space learning. We introduce graph metanetworks that predict properties of neural-network weights and generalize out of distribution to input networks with much larger widths.

graph neural networks metanetworks transferability size generalization mup tensor program

μpscaling small models: Principled warm starts and hyperparameter transfer

Yuxin Ma, Nan Chen, Mateo Díaz, Soufiane Hayou, Dmitriy Kunisky, Soledad Villar 

ICML 2026.

We propose a theoretically grounded model-upscaling method that uses a trained smaller neural network to initialize a larger one. The algorithm is based on dynamic equivalences between models of different widths and the μP infinite-width scaling limit, enabling zero-shot hyperparameter transfer.

model upscaling hyperparameter transfer mup tensor program training dynamics infinite-width limit

On transferring transferability: Towards a theory for size generalization

Eitan Levin*, Yuxin Ma*, Mateo Díaz, Soledad Villar 

NeurIPS 2025 (Spotlight).

We introduce a mathematical framework for understanding when machine learning models generalize across input dimensions, such as graph neural networks across graph sizes, Deep Sets across set sizes, and transformers across sequence lengths.

transferability size generalization graph neural networks equivariant machine learning any-dimensional learning

Nonlinear Laplacians: Tunable principal component analysis under directional prior information

Yuxin Ma, Dmitriy Kunisky 

NeurIPS 2025 (Spotlight).

We study a new class of spectral algorithms for low-rank estimation based on tunable nonlinear deformations of the observed matrix. The deformation can be chosen by black-box optimization or learned from data using neural networks.

principal component analysis random matrix theory spiked matrix models low-rank estimation