Skip to content
Sıfırıncı Dakika
Latest

AWS introduces search agent fine-tuning with multi-turn RL on SageMaker AI

AIUpdated: 2 min read
AWS Machine Learning Blog

In brief

AWS announced fine-tuning of search agents with multi-turn reinforcement learning (MTRL) on Amazon SageMaker AI. The approach optimizes the agent across the full multi-turn trajectory rather than single responses, aiming to combine a small model's speed and cost with the reliability that would otherwise require a frontier model. AWS shared the results it observed in retrieval quality and reliabili

Why search agents are hard

Search agents powered by large language models decide on their own what to search for, which retrieval strategy to use, and when to stop, instead of requiring users to craft the perfect query. They refine these decisions across multiple rounds based on what they have already retrieved.

The problem: no base model arrives knowing your tools and environment. Prompt a small model and dependable multi-turn behavior is rare; prompt a frontier model and it often works, but you pay in latency and cost.

Fine-tuning and MTRL

According to AWS, fine-tuning offers a third path: teaching a small model your tools and environment directly. The result is a small model's speed and cost with the reliability that would otherwise require a frontier model.

Traditional approaches fall short. Supervised fine-tuning (SFT) depends on expert demonstrations of ideal multi-turn trajectories, which are costly to collect and usually don't exist for a given setup. Single-turn reinforcement learning scores one response at a time and misses the dependencies between interdependent multi-turn decisions.

Multi-turn reinforcement learning (MTRL) optimizes the agent across the full multi-turn trajectory. The reward signal only needs to reflect whether the final outcome was good.

What SageMaker AI MTRL offers

  • Modular agent-environment interface: define custom rewards, custom tool loops, and multi-turn conversation shapes.
  • Serverless execution: production-scale agentic RL at per-token pricing without provisioning GPU clusters.
  • Asynchronous rollout and trajectory collection: generation and gradient updates run in parallel.
  • Native algorithm library: PPO, CISPO, and importance-sampling losses, paired with group-based advantage estimators such as GRPO, GRPO pass@k, and RLOO.
  • Resumable training: split long runs across multiple jobs.
  • Trajectory and reward observability: inspect what the agent did turn by turn in MLflow.
  • Evaluation jobs: report reward, pass@k, and trajectory metrics before deployment.

AWS frames MTRL as a natural fit for search agents: a clear reward signal (retrieval quality), a multi-turn interaction loop, and a well-defined environment (the search tools).

Why it matters

Teams building enterprise search agents have often been stuck paying for frontier models to get reliable results. AWS's method trains a small model across the whole conversation to reach similar reliability at a fraction of the cost, a directly usable option for teams already on SageMaker.

Sources

  1. AWS Machine Learning BlogFirst reported byPrimary source
    Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

Related stories