Eigen RadarAI
Analysis

LinkedIn trains its job-search ranking model eight times faster via multi-teacher distillation

LinkedIn built a job-search ranker containing 0.6 billion parameters, using supervisory signals from several larger models. It mixed teacher answers prepared in advance with answers generated during training, reducing total training time eightfold. The company also reported large gains in ranking throughput and NDCG@10 on its own system.

Artificial Intelligence··Evening
Several large teacher volumes send guidance into a compact ranking core in a conceptual computational scene.

A smaller ranker from several teachers

LinkedIn trained a job-search ranking model with 0.6 billion parameters by transferring signals from several larger teacher models. The method combines teacher answers prepared before training with additional answers generated while training continues. LinkedIn reports that this mixed approach shortened the full training process by a factor of eight.[1]

Several changes supplied the speedup

The gain did not come from one component. A LiGer kernel let each batch hold twice as much work. Training across multiple machines added a 3.5-fold gain, FSDP2 improved the rate by another 20 percent, and moving the workload to H200 clusters contributed a further 30 percent. These figures describe LinkedIn’s own training stack.[1]

Ranking rate and measured quality rose

For each graphics processor, the system’s ranking rate increased from 290 items a second to more than 2,000. Its NDCG@10 score, a ranking-quality measure, moved from 0.7583 to 0.9432. The report presents these as results from LinkedIn’s job-search system; it does not provide an independent reproduction on another service or data set.[1]

References

  1. News sourceInfoQLinkedIn trains its job-search ranker eight times faster with multi-teacher distillation↩1↩2↩3