Back to blog
AI Agents

Tool Calling in AI Agents: How Models Invoke Real Functions

Model proposes the call; the application decides whether to run it

Typed contracts, permission checks, retries and read-back

Original TOOL-GATE-8 execution contract and failure routes
How a model proposes functions while the application controls execution
Primary nodeScoped tool execution
Routing modeTOOL-GATE-8
StatusPUBLISHED
Tool call proposal passes through schema and permission gates to execution, verification and failure routes
TOOL_GATE_8_V01: route a model proposal through a policy gate to a verified result.
TERMINAL_PREVIEW.LOG
$ call --contract TOOL-GATE-8
> propose: tool / typed arguments
> gate: identity / tenant / policy
> execute: bounded adapter / idempotency
> verify: read-back / receipt / route
Tool execution with explicit boundaries

Tool calling is a contract between a model, an application and an external system. The model proposes a function name and arguments. The application checks the schema, identity, permissions and risk, executes an allowed operation, then returns a bounded result. The model never receives direct access to the database, CRM or mail service. That boundary matters more than the call syntax.

This guide follows a call from a user request to a verified result. The example is synthetic and illustrates an architecture, not measured customer performance. If the choice of an agent is still open, start with the task-first decision tree. Broader implementation questions belong on the AI specialist Armenia page.

Problem and requirements

Consider: “Find the current supplier record and prepare an address change.” The system needs a read, source check and draft. A CRM write is a separate operation requiring approval. A model saying “updated” proves nothing: the destination system must acknowledge the operation and be read again. Define the owner, tenant, object ID, time budget, permitted tools and acceptance criterion before adding tool access.

Functions are useful when the answer needs an external fact or action. Plain language must not conceal a missing call. Nor should a write function be exposed to the model merely because an API offers it. Give each stage the smallest permitted tool catalogue.

Architecture: proposal, gateway, execution, receipt

An authenticated request reaches the orchestrator. It supplies a task and limited tool schemas. The model proposes a call. A gateway treats that proposal as untrusted input: it validates name, strict schema, size, tenant, user permissions and action policy. Only then does an adapter contact the external service. The result is normalized and redacted; a receipt is stored; the model sees only the needed fragment.

NodeContractCheck
ModelTool name and JSON argumentsSchema, allowed values, step budget
Policy gatewayIdentity, tenant, action, objectPermissions independent of model text
AdapterTimeout, idempotency key, result codeRetry, recovery, unknown outcome
State storeRequest ID, versions, route, receiptResume without duplicate writes
Result verifierRead-back from destinationWas the action actually accepted?

The diagram is a design aid, not an autonomy guarantee. It makes failure boundaries inspectable. AI automation covers process integration; prompt engineering can shape output, but cannot replace server-side authorization.

Original example: TOOL-GATE-8

The minimal example offers supplier.lookup as a read-only operation and supplier.draft_change for a draft. The model cannot call supplier.commit_change before human approval. All IDs and values are illustrative; this is contract pseudocode rather than a ready-made SDK.

ts
type ProposedCall = { name: string; args: unknown; requestId: string };
async function execute(call: ProposedCall, ctx: Context) {
  const tool = registry.get(call.name);
  if (!tool) return deny("unknown_tool");
  const args = tool.schema.parse(call.args);
  if (!policy.allows(ctx.user, ctx.tenant, call.name, args)) return deny("forbidden");
  if (tool.risk === "write" && !ctx.approvalFor(call.requestId)) return hold("approval_required");
  const key = `${ctx.tenant}:${call.requestId}:${call.name}`;
  const result = await withTimeout(tool.run(args, { key, tenant: ctx.tenant }), tool.timeoutMs);
  const checked = tool.risk === "write" ? await tool.readBack(args) : result;
  await receipts.save({ key, tool: call.name, route: "verified", resultRef: checked.ref });
  return redact(checked);
}

A real implementation returns a structured rejection on parse errors instead of retrying with wider permissions. An idempotency key prevents a network retry from creating a second change. A timeout may leave an unknown outcome: query status with the same key before retrying. Logs should retain references and reason codes, not secret payloads.

Key components and boundaries

Tool schema. Version the name, purpose, fields, allowed values and result type. Validate server-side even when a model produces structured output. Reject empty strings, wrong types, unexpected tenants and oversized payloads before the adapter.

Permissions and approval. The user session determines data scope. Instructions embedded in a document or tool result do not expand it. Reading, drafting and committing are separate capabilities. Never put a service credential in a model prompt or treat “please do not write” as an access control.

State. Persist request ID, chosen tool, schema version, idempotency key, rejection reason and terminal route. Long tasks need resume logic that cannot repeat a completed write.

Result. Normalize upstream errors. Distinguish denied, invalid, timeout, unknown_outcome, completed and verified. After a write, read the object from the destination or use an equivalent authoritative confirmation. A successful HTTP response alone does not prove a business outcome.

Failure modes

  1. Unknown tool. The model proposes an absent name. Reject explicitly and limit new proposals.
  2. Invalid arguments. Allow declared fields only; model text cannot replace server-bound object or tenant scope.
  3. Prompt injection in read data. Supplier text may contain instructions. Treat it as data, not policy or approval.
  4. Timeout after write. The outcome is unknown. Query by idempotency key and read back before retrying.
  5. Stale version. A version change between read and write must produce a conflict and new review.
  6. False completion. The model claims success without a receipt. The UI should show application-verified status.
  7. Unbounded loop. Step, time and cost budgets stop the cycle independently of model intent.

Testing and production checklist

Build fixtures for a valid call, unknown name, wrong type, cross-tenant object, missing permission, injected tool result, timeout before and after execution, duplicated request ID, version conflict and failed read-back. Declare the expected route and forbidden side effect for each. Test adapters and the policy gateway independently, then run the full receipt path.

Before release, record schema and policy versions, tool owner, execution budget, stop and recovery procedure, redaction rules and retention. Exercise staging with test credentials. Pilot a narrow task set; expand only after failure review and human acceptance. An architecture review may show that a deterministic workflow is sufficient.

For a concrete tool catalogue, permissions and result-verification design, request an architecture review. Bring your system map, permitted actions and one failed-case example.

CODE_BLOCK.TXT
require(schema.valid && policy.allows(user, tenant, tool));
if (tool.isWrite && !approval.valid) return hold;
result = await executeOnce(idempotencyKey);
if (result.unknown) return reconcileBeforeRetry;
return verifyByReadBack(result);