🧬 "More data is always better" is a machine learning truism that may not hold up in genomic medicine.
Title: Transfer learning for genomic prediction in underrepresented populations
URL:
Polygenic risk scores (PRS) predict disease risk, but they've mostly been built on European cohorts, so accuracy drops when applied to other populations. This study varied sample sizes for both UK Biobank (European) and Biobank Japan across eight clinical traits to test how far transfer learning can help.
Highlight 1 📉 Bigger isn't always better
Transfer learning from European data only helps predictions for the Japanese population up to about 15,000 target samples. Beyond that crossover point, models trained solely on local data actually perform better.
Highlight 2 🧬 The "shelf life" depends on the trait
Genetically conserved traits like BMI keep benefiting from European data up to 25,000-40,000+ samples, while population-specific traits like lipid levels lose that benefit much sooner.
Highlight 3 ⚖️ The best method depends on sample size too
The cross-population method PRS-CSx underperforms simpler approaches at small sample sizes but catches up as sample size approaches 100,000. There's no universally best method.
The takeaway: equitable clinical genomics needs both larger local biobanks and modeling strategies matched to each trait's genetic architecture and sample size.
#
Genomics# #
TransferLearning#