Perplexity.AI reduces tool-call failures by 21% with new model training

Summary

New research reveals a significant advancement in AI training, where a computer model improved its performance by learning from its own errors through hint-guided self-distillation. In a live A/B test, a later trained checkpoint successfully reduced tool-call failures by 21.2% compared to an earlier version. This process not only allows the model to learn from real-world sessions that highlight previously unseen failures but also ensures that training avoids personally identifiable information, thereby supporting safer interactions in production environments. By utilizing corrective hints during training, the model was able to avoid original errors in 93.7% of cases, indicating a marked enhancement in its learning process.

Analysis

GLM 5.2: GLM 5.2 is an advanced AI language model evaluated in post-training experiments. The research uses it as the base model for hint-guided self-distillation to align predictions and reduce errors in tool-calling scenarios from real-world sessions. perplexity.ai: perplexity.ai is an AI search and research platform that published this study on improving model performance through real-world data. The company details a pipeline that filters sessions, applies rejection sampling, and uses validated hints to correct model mistakes without introducing hindsight bias. Model Improvement: Post-training pipelines that separate PII and opted-out data support safer learning from live user interactions in production environments. AI Training Methods: Hint-guided self-distillation enables models to internalize corrections from both successful and failed sessions while preserving original decision context.

Categories

machine_learningaitech
View Original Tweet