NASA’s new dark energy space telescope can also detect killer asteroids
technologyreview.com·1h ago
TL;DR
In low-resource multi-label classification, there is uncertainty about when synthetic patent data can enhance performance. The study utilized six open-source large language models (LLMs) to generate synthetic data for classifying 64 patent labels.
✦ Why It Matters
Engineers can optimize multi-label classification by strategically combining real and synthetic data to enhance performance.
Key Takeaways
How It Works
The study employs six LLMs, ranging from 3.8B to 12B parameters, to generate synthetic patent data. It tests both full-synthesis generation and paraphrasing methods, assessing their impact on classification performance across various data regimes.
Related