Compare a broad tool and a focused tool using measurable task outcomes.
A tool should help the agent express a useful intent with limited ambiguity. Exposing every internal endpoint can overload selection. Exposing one universal command can hide dangerous choices. The right boundary depends on the task and the service invariants, so compare alternatives with concrete cases.
Start with the user's objective and the decisions that require judgment. If the user asks to find available interview slots, the service can calculate availability deterministically. The model does not need to inspect every calendar event and implement overlap logic in its response. A focused availability tool can return a small set of valid slots with explicit time zones and expiry. The agent then helps the user choose.
Combining reads is usually easier than combining writes. A tool that searches a customer and returns relevant subscription details reduces round trips. A tool that searches, selects, cancels, refunds, and emails in one step hides several independent consequences. If a combined tool is justified, its contract must define preconditions, partial failure, and evidence for every effect. "One tool call" does not mean "one atomic transaction".
Measure the boundary with a fixed task set. Include ambiguity, missing data, stale versions, and unauthorized resources. Count correct outcomes, unintended effects, recovery steps, latency, and response size. Do not score the agent only for matching one expected call sequence. Several safe sequences may produce the same valid outcome. Conversely, a plausible final sentence is insufficient when the service state is wrong.
Worked example
A fictional support system offers either twelve low-level contact tools or two focused tools: search contacts and draft a reply for a selected contact. Twenty tasks contain duplicate names, old addresses, and missing consent. In the teaching results, the low-level design resolves 14 tasks correctly with 120 calls. The focused design resolves 18 with 56 calls. Both have zero sends because the tools only create drafts.
The result supports trying the focused interface for this task distribution. It does not prove that fewer tools always win. Review the two failures. Both involve organizations with shared inboxes, so the next revision adds organization ID and address status to search results. Adding automatic send would change the safety boundary and require a separate evaluation.
Exercise and solution
Design a tool set for "Find three suitable rooms for a 20-person meeting next Tuesday". The system has room capacity, equipment, location, and calendar records. The user has not authorized a booking.
A good solution provides a search tool with capacity, date interval, location, and required equipment. Its result contains room IDs, available intervals, time zones, and freshness. It offers a separate booking tool only after a room and time are selected under the application's approval policy. Award one point each for deterministic filtering, explicit time, read/write separation, and a test involving a stale room slot. Do not award credit for a tool that books the first result while claiming it merely searched.
Compare two inspectable interfaces
The following fictional interfaces support finding a contact. They are proposals to evaluate, not a recommendation to expose unrestricted directory data.
Interface B reduces irrelevant output and exposes the disambiguating fields in one response. It also restricts result volume. However, it can fail if its search semantics hide exact matches or rank an old address first without marking it. Tool simplicity must be evaluated against the actual task, including ambiguity and stale data.
Use this original task matrix to compare both interfaces.
Case
Required result
Failure to detect
Two people share a name
Ask or use supplied organization
Selecting the first name match
Address is inactive
Report status and find valid evidence
Sending to the old address
No authorized match
Return no accessible match
Searching another tenant
Many matches
Paginate or narrow query
Pretending the first page is complete
Exact ID supplied
Resolve that authorized ID
Replacing it with a similar name
The matrix tests mechanisms that a single happy-path example misses. It also separates search quality from send permission. Both interfaces can remain read-only during evaluation. The service still checks which contacts the user may access.
A second worked comparison
Suppose the original twenty-task comparison showed fewer calls with Interface B. A reviewer notices that all tasks used exact names and contained fewer than ten contacts. That distribution favors the new interface and does not test pagination or ambiguous partial queries.
Add ten held-out tasks covering those conditions. In the teaching result, B resolves six and A resolves eight. The correct response is to inspect why B fails, not discard the new tasks to preserve its earlier win. Perhaps B truncates results without a continuation marker. Adding an explicit cursor and a "more matches exist" field addresses that failure hypothesis. The revised interface then needs evaluation on both the earlier tasks and a fresh held-out set.
This is an example of test-set reuse risk. If the same ten failures guide every revision, they become development cases. Keep some independent tasks for final estimation. Record which set served which role so the measured gain is not presented as untouched evidence.
Misconceptions to reject
"Fewer calls means a better tool" ignores correctness, disclosure, and hidden effects. A tool can reduce calls by returning every confidential record or silently taking an action. Neither is a valid improvement.
"One tool per user goal is always right" ignores goals with several independent consequences. A combined search-refund-email tool may make recovery and approval harder because its effects have different preconditions and providers.
Transfer exercise
Design a contact-search result for three people named Mira at two organizations, one with an inactive address. The user's request names the organization but not the address. A model answer returns stable contact IDs, display names, organization IDs, address status, and enough match evidence to choose or ask a focused question. It excludes private fields irrelevant to selection. If two authorized active contacts still match, it reports the ambiguity instead of guessing.
Score one point for each disambiguating field, one for minimum necessary data, and one for preserving an unresolved ambiguity. Explain how the result would change if the user supplied an exact contact ID. The service should resolve that ID under authorization, while the model should not substitute a nearby name match.
Interview probe
Original practice: Would one tool per endpoint make an agent easier to build? A strong answer uses task-level ambiguity and effect boundaries to choose granularity. Follow up with a multi-effect transaction that partly fails. A weak answer says that fewer tools are always better without a task distribution or measurement.
A tool returns ten matches but no continuation marker. Which interpretation is safe?
AThe result may be truncated, so completeness cannot be assumed from the ten rows.BRepeating the identical request necessarily returns the next ten matches.CFiltering these ten rows locally is equivalent to filtering the full directory.DThe most relevant identity must be present because the result limit is ten.
When is a combined search-refund-email tool most concerning?
AWhen a read-only search response includes the fields needed to disambiguate the customer.BWhen related read operations are combined under the same access policy.CWhen stable IDs and version preconditions are required for the selected resource.DWhen it uses a shared authorization precheck but does not recheck separate effect permissions or report partial completion.
A revision was tuned repeatedly on ten failure cases. How should those cases be classified for final evaluation?
AUntouched test cases.BDevelopment cases, with fresh held-out evidence still needed.CInvalid cases to delete.DProof of general improvement once all pass.
Two active authorized contacts still match the supplied organization and name. What should the tool-assisted agent do?
ASend to both contacts so at least the intended person receives the message.BChoose the first result because both are authorized.CExpose the remaining ambiguity and ask for a distinguishing detail.DChoose the contact with the most recent record update even though the user supplied no such preference.
Without looking at the worked solution, explain how you would compare two tool interfaces using ambiguous, stale, and paginated cases. Name the service evidence you need. Rate confidence from 1 to 5 and identify the boundary you still cannot justify.
Not yetGetting thereConfident
Wrap-up
Use the smallest interface that expresses intent without hiding consequential decisions. Evaluate effects and recovery, not attractive call traces.