Developing Training for New Technologies

Explore top LinkedIn content from expert professionals.

  • View profile for Andrew Ng
    Andrew Ng Andrew Ng is an Influencer

    DeepLearning.AI, AI Fund and AI Aspire

    2,586,732 followers

    Inexpensive token generation and agentic workflows for LLMs open up new possibilities for training LLMs on synthetic data. Pretraining an LLM on its own directly generated responses to prompts doesn't help. But if an agentic workflow implemented with the LLM results in higher quality output than the LLM can generate directly, then training on that output becomes potentially useful. Just as humans can learn from their own thinking, perhaps LLMs can, too. Imagine a math student learning to write mathematical proofs. By solving a few problems — even without external input — they can reflect on what works and learn to generate better proofs. LLM training involves (i) pretraining (learning from unlabeled text data to predict the next work) followed by (ii) instruction fine-tuning (learning to follow instructions) and (iii) RLHF/DPO to align to human values. Step (i) requires orders of magnitude more data than the others. For example, Llama 3 was pretrained on over 15 trillion tokens. LLM developers are still hungry for more data. Where can we get more text to train on? Many developers train smaller models on the output of larger models, so a smaller model learns to mimic a larger model’s behavior on a particular task. But an LLM can’t learn much by training on data it generated directly. Indeed, training a model repeatedly on the output of an earlier version of itself can result in model collapse. But, an LLM wrapped in an agentic workflow can produce higher-quality output than it can generate directly. This output might be useful as pretraining data. Efforts like these have precedents: - When using reinforcement learning to play a game like chess, a model might learn a function that evaluates board positions. If we apply game tree search along with a low-accuracy evaluation function, the model can come up with more accurate evaluations. Then we can train that evaluation function to mimic these more accurate values. - During alignment, Anthropic’s constitutional AI uses RLAIF (RL from AI Feedback) to judge LLM output quality, substituting feedback generated by an AI model for human feedback. A significant barrier to using agentic workflows to produce LLM training data is the cost of generating tokens. Say we want to generate 1 trillion tokens to extend a pre-existing dataset. At current retail prices, 1 trillion tokens from GPT-4-turbo ($30 per million output tokens), Claude 3 Opus ($75), Gemini 1.5 Pro ($21), and Llama-3-70B on Groq ($0.79) would cost, respectively, $30M, $75M, $21M and $790K. Of course, an agentic workflow would require generating more than one token per final output token. But budgets for training cutting-edge LLMs easily surpass $100M, so spending a few million dollars more for data to boost performance is feasible. That’s why agentic workflows might opening up new opportunities for high-quality synthetic data generation. [Original text: https://lnkd.in/gFF2AsZ9 ]

  • View profile for Andreas Horn

    VP of AI + Growth @ BLP || Speaker | Lecturer | Advisor | Author

    252,724 followers

    OpenAI 𝗷𝘂𝘀𝘁 𝗿𝗲𝗹𝗲𝗮𝘀𝗲𝗱 𝘁𝗵𝗲𝗶𝗿 𝗹𝗲𝗮𝗱𝗲𝗿𝘀𝗵𝗶𝗽 𝗽𝗹𝗮𝘆𝗯𝗼𝗼𝗸 𝗼𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗮𝘆 𝗮𝗵𝗲𝗮𝗱 𝗶𝗻 𝘁𝗵𝗲 𝗮𝗴𝗲 𝗼𝗳 𝗔𝗜 — 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗺𝗲𝘀𝘀𝗮𝗴𝗲 𝗶𝘀 𝗯𝗹𝘂𝗻𝘁: ⬇️ "𝖳𝗁𝖾 𝖼𝗈𝗆𝗉𝖺𝗇𝗂𝖾𝗌 𝗍𝗁𝖺𝗍 𝗐𝗂𝗅𝗅 𝗍𝗁𝗋𝗂𝗏𝖾 𝖺𝗋𝖾 𝗍𝗁𝖾 𝗈𝗇𝖾𝗌 𝗍𝗁𝖺𝗍 𝗍𝗋𝖾𝖺𝗍 𝖠𝖨 𝗇𝗈𝗍 𝗃𝗎𝗌𝗍 𝖺𝗌 𝖺 𝗍𝗈𝗈𝗅, 𝖻𝗎𝗍 𝖺𝗌 𝖺 𝗇𝖾𝗐 𝗐𝖺𝗒 𝗈𝖿 𝗐𝗈𝗋𝗄𝗂𝗇𝗀." AI adoption is moving faster than most leaders ever imagined. Staying ahead is about creating the right conditions for your people and teams to adapt with confidence. The report distills lessons from leaders at Moderna, Notion, BBVA et.al. into practical steps that any company can act on now. 𝗛𝗲𝗿𝗲 𝗮𝗿𝗲 𝘁𝗵𝗲 𝟱 𝗲𝘀𝘀𝗲𝗻𝘁𝗶𝗮𝗹𝘀 𝗢𝗽𝗲𝗻𝗔𝗜 𝗿𝗲𝗰𝗼𝗺𝗺𝗲𝗻𝗱𝘀: ⬇️ 1. 𝗔𝗹𝗶𝗴𝗻  → Start with clarity of purpose. Show your teams why AI matters, set company-wide goals, and role-model adoption at every level. Alignment builds trust and helps employees connect their daily work to your broader AI strategy 2. 𝗔𝗰𝘁𝗶𝘃𝗮𝘁𝗲  → Training > talk. Make learning real and practical. Invest in structured training, create AI champions, and give people room to experiment. When employees see AI as part of their growth and success, adoption becomes natural. 3. 𝗔𝗺𝗽𝗹𝗶𝗳𝘆  → Don’t let wins live in silos. Share success stories widely, build knowledge hubs, and create active communities so everyone can learn from what’s working. Momentum spreads fastest when people see peers succeeding. 4. 𝗔𝗰𝗰𝗲𝗹𝗲𝗿𝗮𝘁𝗲  → Remove friction. Make it easy for teams to access tools, submit ideas, and move projects from pilot to production. Empower decision-making and reward teams who push ideas forward. 5. 𝗚𝗼𝘃𝗲𝗿𝗻  → Balance speed with responsibility. Clear, lightweight guidelines ensure progress without unnecessary bottlenecks. When governance is practical and evolving, it protects the business while keeping innovation alive. This is a surprisingly good read! The recommendations are sharp because they cut right into the real bottlenecks of AI adoption: not the models, but people, processes, and governance. The full guide has only 15 pages: https://lnkd.in/defvM4cj 𝗣.𝗦. 𝗜 𝗿𝗲𝗰𝗲𝗻𝘁𝗹𝘆 𝗹𝗮𝘂𝗻𝗰𝗵𝗲𝗱 𝗮 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿 𝘄𝗵𝗲𝗿𝗲 𝗜 𝘀𝗵𝗮𝗿𝗲 𝘁𝗵𝗲 𝗯𝗲𝘀𝘁 𝘄𝗲𝗲𝗸𝗹𝘆 𝗱𝗿𝗼𝗽𝘀 𝗼𝗻 𝗔𝗜 𝗮𝗴𝗲𝗻𝘁𝘀, 𝗲𝗺𝗲𝗿𝗴𝗶𝗻𝗴 𝘄𝗼𝗿𝗸𝗳𝗹𝗼𝘄𝘀, 𝗮𝗻𝗱 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗮𝘆 𝗮𝗵𝗲𝗮𝗱 𝘄𝗵𝗶𝗹𝗲 𝗼𝘁𝗵𝗲𝗿𝘀 𝘄𝗮𝘁𝗰𝗵 𝗳𝗿𝗼𝗺 𝘁𝗵𝗲 𝘀𝗶𝗱𝗲𝗹𝗶𝗻𝗲𝘀. 𝗜𝘁’𝘀 𝗳𝗿𝗲𝗲 — 𝗮𝗻𝗱 𝗮𝗹𝗿𝗲𝗮𝗱𝘆 𝗿𝗲𝗮𝗱 𝗯𝘆 𝟮𝟬,𝟬𝟬𝟬+ 𝗽𝗲𝗼𝗽𝗹𝗲. 𝗝𝗼𝗶𝗻 𝘁𝗵𝗲𝗺 𝗵𝗲𝗿𝗲: https://lnkd.in/dbf74Y9E

  • View profile for Kyle Poyar

    Founder, Growth Unhinged | GTM & Monetization Newsletter

    112,851 followers

    AI products like Cursor, Bolt and Replit are shattering growth records not because they're "AI agents". Or because they've got impossibly small teams (although that's cool to see 👀). It's because they've mastered the user experience around AI, somehow balancing pro-like capabilities with B2C-like UI. This is product-led growth on steroids. Yaakov Carno tried the most viral AI products he could get his hands on. Here are the surprising patterns he found: (Don't miss the full breakdown in today's bonus Growth Unhinged: https://lnkd.in/ehk3rUTa) 1. Their AI doesn't feel like a black box. Pro-tips from the best: - Show step-by-step visibility into AI processes - Let users ask, “Why did AI do that?” - Use visual explanations to build trust. 2. Users don’t need better AI—they need better ways to talk to it. Pro-tips from the best: - Offer pre-built prompt templates to guide users. - Provide multiple interaction modes (guided, manual, hybrid). - Let AI suggest better inputs ("enhance prompt") before executing an action. 3. The AI works with you, not just for you. Pro-tips from the best: - Design AI tools to be interactive, not just output-driven. - Provide different modes for different types of collaboration. - Let users refine and iterate on AI results easily. 4. Let users see (& edit) the outcome before it's irreversible. Pro-tips from the best: - Allow users to test AI features before full commitment (many let you use it without even creating an account). - Provide preview or undo options before executing AI changes. - Offer exploratory onboarding experiences to build trust. 5. The AI weaves into your workflow, it doesn't interrupt it. Pro-tips from the best: - Provide simple accept/reject mechanisms for AI suggestions. - Design seamless transitions between AI interactions. - Prioritize the user’s context to avoid workflow disruptions. -- The TL;DR: Having "AI" isn’t the differentiator anymore—great UX is. Pardon the Sunday interruption & hope you enjoyed this post as much as I did 🙏 #ai #genai #ux #plg

  • View profile for Aishwarya Srinivasan
    Aishwarya Srinivasan Aishwarya Srinivasan is an Influencer
    646,637 followers

    If you are building AI agents or learning about them, then you should keep these best practices in mind 👇 Building agentic systems isn’t just about chaining prompts anymore, it’s about designing robust, interpretable, and production-grade systems that interact with tools, humans, and other agents in complex environments. Here are 10 essential design principles you need to know: ➡️ Modular Architectures Separate planning, reasoning, perception, and actuation. This makes your agents more interpretable and easier to debug. Think planner-executor separation in LangGraph or CogAgent-style designs. ➡️ Tool-Use APIs via MCP or Open Function Calling Adopt the Model Context Protocol (MCP) or OpenAI’s Function Calling to interface safely with external tools. These standard interfaces provide strong typing, parameter validation, and consistent execution behavior. ➡️ Long-Term & Working Memory Memory is non-optional for non-trivial agents. Use hybrid memory stacks, vector search tools like MemGPT or Marqo for retrieval, combined with structured memory systems like LlamaIndex agents for factual consistency. ➡️ Reflection & Self-Critique Loops Implement agent self-evaluation using ReAct, Reflexion, or emerging techniques like Voyager-style curriculum refinement. Reflection improves reasoning and helps correct hallucinated chains of thought. ➡️ Planning with Hierarchies Use hierarchical planning: a high-level planner for task decomposition and a low-level executor to interact with tools. This improves reusability and modularity, especially in multi-step or multi-modal workflows. ➡️ Multi-Agent Collaboration Use protocols like AutoGen, A2A, or ChatDev to support agent-to-agent negotiation, subtask allocation, and cooperative planning. This is foundational for open-ended workflows and enterprise-scale orchestration. ➡️ Simulation + Eval Harnesses Always test in simulation. Use benchmarks like ToolBench, SWE-agent, or AgentBoard to validate agent performance before production. This minimizes surprises and surfaces regressions early. ➡️ Safety & Alignment Layers Don’t ship agents without guardrails. Use tools like Llama Guard v4, Prompt Shield, and role-based access controls. Add structured rate-limiting to prevent overuse or sensitive tool invocation. ➡️ Cost-Aware Agent Execution Implement token budgeting, step count tracking, and execution metrics. Especially in multi-agent settings, costs can grow exponentially if unbounded. ➡️ Human-in-the-Loop Orchestration Always have an escalation path. Add override triggers, fallback LLMs, or route to human-in-the-loop for edge cases and critical decision points. This protects quality and trust. PS: If you are interested to learn more about AI Agents and MCP, join the hands-on workshop, I am hosting on 31st May: https://lnkd.in/dWyiN89z If you found this insightful, share this with your network ♻️ Follow me (Aishwarya Srinivasan) for more AI insights and educational content.

  • View profile for Vitaly Friedman
    Vitaly Friedman Vitaly Friedman is an Influencer

    Practical insights for better UX • Running “Measure UX” and “Design Patterns For AI” • Founder of SmashingMag • Speaker • Loves writing, checklists and running workshops on UX. 🍣

    231,811 followers

    🎢 Onboarding UX Playbook (+ Decision Trees). Practical techniques for better onboarding UX, design patterns, kits and Figma templates — on mobile and desktop. 🚫 Users often skip tutorials/walkthroughs entirely. 🚫 Never block the UI with full-page onboarding modals. 🚫 Avoid long multi-step tutorials with 5+ steps. ✅ Ask customers what goals they are trying to achieve. ✅ Allow users to hide walkthroughs and restore them later. ✅ Focus on bringing users to first success moments fast. ✅ Structure your onboarding suggestions in bite-sized chunks. ✅ Explain features when users slow down or make mistakes. ✅ Show features when users lose time with repetitive tasks. ✅ Prevent failure with an early warning system for new users. ✅ Collapsible checklists work well for onboarding. ✅ Personalized onboarding works even better. ✅ Design sets of filters, templates and empty states. ✅ Show starter kits based on user’s profile and interests. ✅ Consider short video guides and email drip campaigns. Good onboarding can’t be generic. It has to be relevant and valuable. Define your user segments first. Design a set of presets to help them get to success moments faster. Think of the questions you need to ask to customize their experience. Think about filters and presets they might need. Onboarding tutorials often appear once and get instantly dismissed, nowhere to be found again. Allow users to find them when they need it. Bring them up when users slow down or make mistakes. And test the discoverability of your features continuously. If a feature is obvious, you might not need to explain it at all. And if it isn’t, perhaps onboarding won’t solve this problem either. Useful resources: How to Choose Onboarding Methods and Components, by NewsKit 👍 Methods: https://lnkd.in/eWn5FPWA Decision Tree: https://lnkd.in/e8TmMDFf Design Patterns: https://lnkd.in/ed7HjzkW Onboarding UX Playbook, by Eleana Gkogka https://lnkd.in/edcDfMFG Complete Onboarding UX Guide (free eBook), by Intercom https://lnkd.in/eAxT6ZM4 User Onboarding Best Practices, by Taras Bakusevych https://lnkd.in/eRwr2tEc Guide to Onboarding, by Phil Byrne https://lnkd.in/esEavgw7 How Spotify Organizes Onboarding in Figma, by Barton Smith, Cliona O'Sullivan https://lnkd.in/ei434tqq Mobile Onboarding Wireframe Flows (Figma template) https://lnkd.in/ekhzWFJz UX Onboarding Patterns, by Eve Weinberg https://lnkd.in/e7_M4kDv #ux #design

  • View profile for Brij Kishore Pandey
    Brij Kishore Pandey Brij Kishore Pandey is an Influencer

    AI Architect & AI Engineer | Building Agentic Systems & Scalable AI Solutions

    736,302 followers

    Training a Large Language Model (LLM) involves more than just scaling up data and compute. It requires a disciplined approach across multiple layers of the ML lifecycle to ensure performance, efficiency, safety, and adaptability. This visual framework outlines eight critical pillars necessary for successful LLM training, each with a defined workflow to guide implementation: 𝟭. 𝗛𝗶𝗴𝗵-𝗤𝘂𝗮𝗹𝗶𝘁𝘆 𝗗𝗮𝘁𝗮 𝗖𝘂𝗿𝗮𝘁𝗶𝗼𝗻: Use diverse, clean, and domain-relevant datasets. Deduplicate, normalize, filter low-quality samples, and tokenize effectively before formatting for training. 𝟮. 𝗦𝗰𝗮𝗹𝗮𝗯𝗹𝗲 𝗗𝗮𝘁𝗮 𝗣𝗿𝗲𝗽𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴: Design efficient preprocessing pipelines—tokenization consistency, padding, caching, and batch streaming to GPU must be optimized for scale. 𝟯. 𝗠𝗼𝗱𝗲𝗹 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 𝗗𝗲𝘀𝗶𝗴𝗻: Select architectures based on task requirements. Configure embeddings, attention heads, and regularization, and then conduct mock tests to validate the architectural choices. 𝟰. 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗦𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 and 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻: Ensure convergence using techniques such as FP16 precision, gradient clipping, batch size tuning, and adaptive learning rate scheduling. Loss monitoring and checkpointing are crucial for long-running processes. 𝟱. 𝗖𝗼𝗺𝗽𝘂𝘁𝗲 & 𝗠𝗲𝗺𝗼𝗿𝘆 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻: Leverage distributed training, efficient attention mechanisms, and pipeline parallelism. Profile usage, compress checkpoints, and enable auto-resume for robustness. 𝟲. 𝗘𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻 & 𝗩𝗮𝗹𝗶𝗱𝗮𝘁𝗶𝗼𝗻: Regularly evaluate using defined metrics and baseline comparisons. Test with few-shot prompts, review model outputs, and track performance metrics to prevent drift and overfitting. 𝟳. 𝗘𝘁𝗵𝗶𝗰𝗮𝗹 𝗮𝗻𝗱 𝗦𝗮𝗳𝗲𝘁𝘆 𝗖𝗵𝗲𝗰𝗸𝘀: Mitigate model risks by applying adversarial testing, output filtering, decoding constraints, and incorporating user feedback. Audit results to ensure responsible outputs. 🔸 𝟴. 𝗙𝗶𝗻𝗲-𝗧𝘂𝗻𝗶𝗻𝗴 & 𝗗𝗼𝗺𝗮𝗶𝗻 𝗔𝗱𝗮𝗽𝘁𝗮𝘁𝗶𝗼𝗻: Adapt models for specific domains using techniques like LoRA/PEFT and controlled learning rates. Monitor overfitting, evaluate continuously, and deploy with confidence. These principles form a unified blueprint for building robust, efficient, and production-ready LLMs—whether training from scratch or adapting pre-trained models.

  • View profile for Sahar Mor

    I help researchers and builders make sense of AI | ex-Stripe | aitidbits.ai | Angel Investor

    42,563 followers

    Researchers at UC San Diego and Tsinghua just solved a major challenge in making LLMs reliable for scientific tasks: knowing when to use tools versus solving problems directly. Their method, called Adapting While Learning (AWL), achieves this through a novel two-component training approach: (1) World knowledge distillation - the model learns to solve problems directly by studying tool-generated solutions (2) Tool usage adaptation - the model learns to intelligently switch to tools only for complex problems it can't solve reliably The results are impressive: * 28% improvement in answer accuracy across scientific domains * 14% increase in tool usage precision * Strong performance even with 80% noisy training data * Outperforms GPT-4 and Claude on custom scientific datasets Current approaches either make LLMs over-reliant on tools or prone to hallucinations when solving complex problems. This method mimics how human experts work - first assessing if they can solve a problem directly before deciding to use specialized tools. Paper https://lnkd.in/g37EK3-m — Join thousands of world-class researchers and engineers from Google, Stanford, OpenAI, and Meta staying ahead on AI http://aitidbits.ai

  • View profile for Kuldeep Singh Sidhu

    Senior Data Scientist @ Walmart | BITS Pilani

    17,161 followers

    All the way from Korea, a novel approach called Mentor-KD significantly improves the reasoning abilities of small language models. Mentor-KD introduces an intermediate-sized "mentor" model to augment training data and provide soft labels during knowledge distillation from large language models (LLMs) to smaller models. Broadly, it’s a two-stage process: 1) Fine-tune the mentor on filtered Chain-of-Thought (CoT) annotations from an LLM teacher. 2) Use the mentor to generate additional CoT rationales and soft probability distributions. The student model is then trained using: - CoT rationales from both the teacher and mentor (rationale distillation). - Soft labels from the mentor (soft label distillation). Results show that Mentor-KD consistently outperforms baselines, with up to 5% accuracy gains on some tasks. Mentor-KD is especially effective in low-resource scenarios, achieving comparable performance to baselines while using only 40% of the original training data. This work opens up exciting possibilities for making smaller, more efficient language models better at complex reasoning tasks. What are your thoughts on this approach?

  • View profile for Wes Bush

    Author of Product-Led Growth & The Product-Led Playbook | I’ve been told I make PLG simple but you tell me!

    43,591 followers

    Signed up for 100+ SaaS products in the last 6 months. These are the 8 best examples of AI onboarding I’ve seen this year. Not hype, real AI used to onboard users in seconds. Took a few hours to put the onboarding flows on a Figma board, with some notes covering exactly how these companies use AI to get users to value faster. Here’s how they are using AI to cut time-to-value down to seconds 👇 1. Relay.app (Context > Content) Instead of asking 20 questions, Relay asks for your LinkedIn URL. The AI scans your profile and auto-configures your agents and workspace instantly. 2. Gamma (Execution > Guidance) Gamma doesn't teach you how to use the editor. It asks for a topic and generates a full slide deck for you in seconds. No more relying on "empty states." 3. Figma (Just-in-Time Education) Figma analyzes your behavior in the canvas. If you get stuck or pause too long, the AI suggests the specific plugin or feature you need right in that moment. 4. Zapier (Outcome > Templates) Templates have taken a back seat. Now, a Copilot ingests your desired outcome and builds the workflow for you. It uses your initial app selections to predict exactly which prompts you need first. 5. Notion (Conversational Setup) They replaced the static "welcome wizard" with an active AI chat. It uses natural language to configure your workspace behind the scenes. 6. Miro (Zero-Click Canvas) The first screen is a chatbot asking, "What are we working on?". It builds the board structure for you before you even learn the UI. 7. n8n (Teaching by Showing) The "Try an AI Workflow" option demonstrates a working example first, teaching you how to interact with the agent while giving you a feeling of immediate progress. 8. Instantly.ai (Embedded Support) While the main tour is traditional (tooltips), the real power is hidden inside. As you navigate, AI agents surface to handle complex setup tasks, proving you don't need to be "AI-Native" to be effective. Onboarding is evolving. → From: Teaching users how to use your interface. → To: Teaching AI what the user wants to do. Think I’m exaggerating? Watch your growth rate when competitors can activate users in seconds, while you do it in minutes. I compiled screenshots of all 8 flows into a Figma Board so you can see exactly how they work. I’m also covering how to do AI onboarding in a live workshop with Mickey Alon next week (Jan 28). Comment "AI Onboarding" below and I'll send you the link for both. 👇

  • View profile for Sohrab Rahimi

    Director, AI/ML Lead @ Google

    24,281 followers

    One of the biggest barriers to deploying LLM-based agents in real workflows is their poor performance on long-horizon reasoning. Agents often generate coherent short responses but struggle when a task requires planning, tool use, or multi-step decision-making. The issue is not just accuracy at the end, but the inability to reason through the middle. Without knowing which intermediate steps helped or hurt, agents cannot learn to improve. This makes long-horizon reasoning one of the hardest and most unsolved problems for LLM generalization. It is relatively easy for a model to retrieve a document, answer a factual question, or summarize a short email. It is much harder to solve a billing dispute that requires searching, interpreting policy rules, applying edge cases, and adjusting the recommendation based on prior steps. Today’s agents can generate answers, but they often fail to reflect, backtrack, or reconsider earlier assumptions. A new paper from Google DeepMind and Stanford addresses this gap with a method called SWiRL: Step-Wise Reinforcement Learning. Rather than training a model to get the final answer right, SWiRL trains the model to improve each step in a reasoning chain. It does this by generating synthetic multi-step problem-solving traces, scoring every individual step using a reward model (Gemini 1.5 Pro), and fine-tuning the base model to favor higher-quality intermediate steps. This approach fundamentally changes the way we train reasoning agents. Instead of optimizing for final outcomes, the model is updated based on how good each reasoning step was in context. For example, if the model generates a search query or a math step that is useful, even if the final answer is wrong, that step is rewarded and reinforced. Over time, the agent learns not just to answer, but to reason more reliably. This is a major departure from standard RLHF, which only gives feedback at the end. SWiRL improves performance by 9.2 percent on HotPotQA, 16.9 percent on GSM8K when trained on HotPotQA, and 11 to 15 percent on other multi-hop and math datasets like MuSiQue, BeerQA, and CofCA. It generalizes across domains, works without golden labels, and outperforms both supervised fine-tuning and single-step RL methods. The implications are substantial: we can now train models to reason better by scoring and optimizing their intermediate steps. Better reward models, iterative reflection, tool-assisted reasoning, and trajectory-level training will lead to more robust performance in multi-step tasks. This is not about mere performance improvement. It shows how we can begin to train agents not to mimic outputs, but to improve the quality of their thought process. That’s essential if we want to build agents that work through problems, adapt to new tasks, and operate autonomously in open-ended environments.

Explore categories