Tutorials › Agentic AI › Planning and Reasoning

Agentic AI · Part 4 of 7

Planning and Reasoning

How an agent decides what to do next: reasoning traces, plans, and reasoning models.

On each trip around the loop, the model decides what to do next from the current transcript. That decision is the agent's planning.

Reasoning before acting

The ReAct paper (Yao et al., 2022) has the model produce “both reasoning traces and task-specific actions in an interleaved manner.” The reasoning helps the model plan and adjust, and the actions let it consult outside sources. The Thought, Action, Observation transcript in Tools and Tool Calling follows this pattern.

The paper addresses the hallucination and error propagation seen in chain-of-thought reasoning, where the model reasons without checking anything outside itself, by letting the model interact with a Wikipedia API. It reports results on question answering and fact verification (HotpotQA and Fever) and on two interactive decision-making benchmarks, ALFWorld and WebShop. On those two, ReAct beat imitation and reinforcement learning methods by absolute success rates of 34% and 10%, respectively, from one or two in-context examples.

Planning first

Plan-and-Solve prompting (Wang et al., 2023) targets the missing-step errors of zero-shot chain-of-thought prompting. The model first devises a plan that divides the task into subtasks, then carries them out. On ten datasets with GPT-3, the authors report that it outperforms zero-shot chain-of-thought across all of them. The paper is about prompting for reasoning problems and involves no tools.

Agent designs also differ in who writes the plan. Anthropic's guide describes prompt chaining, where a developer fixes the steps and “each LLM call processes the output of the previous one,” and orchestrator-workers, where a central model “breaks down tasks, delegates them to worker LLMs, and synthesizes their results,” with the subtasks determined at runtime. In the first, the developer wrote the plan. In the second, the model did.

Reasoning models

Some models are trained to think before they answer. OpenAI's documentation describes reasoning models that use internal reasoning tokens before producing a response. The tokens are not visible through the API, but they are billed as output tokens and occupy space in the context window. A reasoning effort setting controls how much the model thinks, and lower values favor speed and lower token use.

Anthropic's documentation describes interleaved thinking: the model thinks between tool calls, reasoning about each tool result before deciding what to do next. Anthropic's guidance says higher thinking budgets enable more comprehensive reasoning, with diminishing returns that depend on the task and at the cost of increased latency.