Model Fusion via Optimal Transport
Explore this paper's citation graph
Summary
This work presents a layer-wise model fusion algorithm for neural networks that utilizes optimal transport to (soft-) align neurons across the models before averaging their associated parameters, and shows that this can successfully yield "one-shot" knowledge transfer between neural networks trained on heterogeneous non-i.i.d. data.
- Type
- preprint
- Published
- 2019-10-12
- Cited by
- 322
- References
- 44
- Access
- Open access
- OpenAlex
- https://openalex.org/W2979556024
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:204512191
Keywords
MNIST database, Computer science, Perceptron, Artificial intelligence, Code (set theory)
References
- Stacked generalization
- Gradient Flows: In Metric Spaces and in the Space of Probability Measures
- A Brief Introduction to Boosting
- Fast Computation of Wasserstein Barycenters
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Barycenters in the Wasserstein Space
- A Relationship Between Arbitrary Positive Matrices and Doubly Stochastic Matrices
- Learning Complex, Extended Sequences Using the Principle of History Compression
- Iterative Bregman Projections for Regularized Transportation Problems
- Sinkhorn Distances: Lightspeed Computation of Optimal Transport
- Regularized Discrete Optimal Transport
- Deep Residual Learning for Image Recognition
- Structured Pruning of Deep Convolutional Neural Networks
- Model compression
- Scale Space and Variational Methods in Computer Vision
- Communication-Efficient Learning of Deep Networks from Decentralized Data
- Axiomatic Attribution for Deep Networks
- An Investigation of How Neural Networks Learn from the Experiences of Peers Through Periodic Weight Averaging
- Influence-Directed Explanations for Deep Convolutional Networks
Cited by
- Federated Learning with Matched Averaging
- WoodFisher: Efficient second-order approximations for model compression
- DJEnsemble: On the Selection of a Disjoint Ensemble of Deep Learning Black-Box Spatio-temporal Models
- Representation Transfer by Optimal Transport
- Heterogeneous Federated Learning
- Optimizing Mode Connectivity via Neuron Alignment
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Multi-Layer Combinatorial Fusion Using Cognitive Diversity
- Fusing Multitask Models by Recursive Least Squares
- On Ensembles, I-Optimality, and Active Learning
- Fed2: Feature-Aligned Federated Learning
- Co-Transport for Class-Incremental Learning
- Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning
- Knowledge Distillation for 6D Pose Estimation by Keypoint Distribution Alignment
- An Optimal Transport Approach to Personalized Federated Learning
- Federated Self-supervised Learning for Heterogeneous Clients
- Renaissance Robot: Optimal Transport Policy Fusion for Learning Diverse Skills
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels
- Git Re-Basin: Merging Models modulo Permutation Symmetries
Related papers
- CryptoDL: Deep Neural Networks over Encrypted Data
- Improving Neural Architecture Search Image Classifiers via Ensemble Learning
- Economical ensembles with hypernetworks
- Better Together: Resnet-50 accuracy with 13x fewer parameters and at 3x speed
- Efficient K-Shot Learning with Regularized Deep Networks
- Snapshot Ensembles: Train 1, get M for free
- Ensemble-Compression: A New Method for Parallel Training of Deep Neural Networks