Hasty Briefsbeta

Bilingual

Knowledge Distillation of Black-Box Large Language Models (2024)

22 days ago
  • Introduces Proxy-KD, a novel method for knowledge distillation (KD) from black-box large language models (LLMs) to smaller models.
  • Addresses the limitation of inaccessible internal states in black-box teachers by using a proxy model to facilitate efficient knowledge transfer.
  • Demonstrates that Proxy-KD enhances KD performance from black-box teachers and surpasses traditional white-box KD techniques.