Why search agents are hard
Search agents powered by large language models decide on their own what to search for, which retrieval strategy to use, and when to stop, instead of requiring users to craft the perfect query. They refine these decisions across multiple rounds based on what they have already retrieved.
The problem: no base model arrives knowing your tools and environment. Prompt a small model and dependable multi-turn behavior is rare; prompt a frontier model and it often works, but you pay in latency and cost.
Fine-tuning and MTRL
According to AWS, fine-tuning offers a third path: teaching a small model your tools and environment directly. The result is a small model's speed and cost with the reliability that would otherwise require a frontier model.
Traditional approaches fall short. Supervised fine-tuning (SFT) depends on expert demonstrations of ideal multi-turn trajectories, which are costly to collect and usually don't exist for a given setup. Single-turn reinforcement learning scores one response at a time and misses the dependencies between interdependent multi-turn decisions.
Multi-turn reinforcement learning (MTRL) optimizes the agent across the full multi-turn trajectory. The reward signal only needs to reflect whether the final outcome was good.
What SageMaker AI MTRL offers
- Modular agent-environment interface: define custom rewards, custom tool loops, and multi-turn conversation shapes.
- Serverless execution: production-scale agentic RL at per-token pricing without provisioning GPU clusters.
- Asynchronous rollout and trajectory collection: generation and gradient updates run in parallel.
- Native algorithm library: PPO, CISPO, and importance-sampling losses, paired with group-based advantage estimators such as GRPO, GRPO pass@k, and RLOO.
- Resumable training: split long runs across multiple jobs.
- Trajectory and reward observability: inspect what the agent did turn by turn in MLflow.
- Evaluation jobs: report reward, pass@k, and trajectory metrics before deployment.
AWS frames MTRL as a natural fit for search agents: a clear reward signal (retrieval quality), a multi-turn interaction loop, and a well-defined environment (the search tools).