Abstract
Adapters have been positioned as a parameter-efficient fine-tuning (PEFT) approach, whereby a minimal number of parameters are added to the model and fine-tuned. However, adapters have not been sufficiently analyzed to understand if PEFT translates to benefits in training/deployment efficiency and maintainability/extensibility. Through extensive experiments on many adapters, tasks, and languages in supervised and cross-lingual zero-shot settings, we clearly show that for Natural Language Understanding (NLU) tasks, the parameter efficiency in adapters does not translate to efficiency gains compared to full fine-tuning of models. More precisely, adapters are relatively expensive to train and have slightly higher deployment latency. Furthermore, the maintainability/extensibility benefits of adapters can be achieved with simpler approaches like multi-task training via full fine-tuning, which also provide relatively faster training times. We, therefore, recommend that for moderately sized models for NLU tasks, practitioners should rely on full fine-tuning or multi-task training rather than using adapters. Our code is available at https://github.com/AI4Bharat/adapter-efficiency.
| Original language | English |
|---|---|
| Title of host publication | CODS-COMAD '24: Proceedings of the 7th Joint International Conference on Data Science & Management of Data (11th ACM IKDD CODS and 29th COMAD) |
| Number of pages | 19 |
| Publisher | Association for Computing Machinery |
| Publication date | 2024 |
| Pages | 136-154 |
| DOIs | |
| Publication status | Published - 2024 |
| Externally published | Yes |
| Event | International Conference on Data Science & Management of Data - Bangalore, India Duration: 4 Jan 2024 → 7 Jan 2024 Conference number: 7 https://dl.acm.org/doi/proceedings/10.1145/3632410 |
Conference
| Conference | International Conference on Data Science & Management of Data |
|---|---|
| Number | 7 |
| Country/Territory | India |
| City | Bangalore |
| Period | 04/01/2024 → 07/01/2024 |
| Internet address |
Keywords
- computational efficiency
- Deep Learning
- Adapter
- Parameter-efficient fine-tuning
- Cross-lingual zero-shot learning
- Multi-task training
- NLU tasks
- Training efficiency
Fingerprint
Dive into the research topics of 'A Comprehensive Analysis of Adapter Efficiency.'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver