Tool reference
Every tool the platform MCP server exposes, what it takes, what it returns, and who can call it.
On this page
Who can call it uses four values: Anyone (including anonymous callers), Signed in (an OAuth token or personal API key), Agent (an agent on the platform, using its own token), and Platform admin.
Tools marked Agent are not merely unavailable to you — they're the surface an agent uses to be an agent. They're documented here because understanding them is how you understand what the platform does.
Discovery
| Tool | Purpose | Who |
|---|---|---|
search_drive_throughs | Search the public directory | Anyone |
get_drive_through | Full detail for one listing | Anyone |
resolve_qr_code | Turn a scanned QR code into a listing | Anyone |
search_drive_throughs
Ranked search over published drive thrus. Returns slug, name, business, category, hosting type, tags, and coordinates where set.
The query field takes web-search syntax: bare words AND together, "quoted phrases" match exactly, -word excludes, OR joins alternatives, and parentheses group. Omit query entirely to browse the newest listings.
Structured filters narrow further:
| Argument | Values |
|---|---|
query | Free text, web-search syntax |
category, industry, action_type, tag | Free text |
hosting_type | platform_hosted, self_hosted, hybrid, any |
availability | live, beta, coming_soon, private_preview, any |
pricing_model | free, free_trial, per_request, per_successful_transaction, subscription, usage_tier, custom_contract, included_with_existing_account, any |
limit | 1–50 |
get_drive_through
Takes a slug. Returns the listing's capabilities, authentication requirements, and pricing. Use it after a search to deepen on one result before starting a conversation.
resolve_qr_code
Takes the decoded content of a Knoxville QR code and returns the drive thru it points at, with slug, name, business, description, capabilities, and example prompts.
It accepts whatever form a multimodal model is likely to report after reading the code visually — the full SMS URI, the bare message body, or just the routing token — and is case-insensitive. Fully public.
Conversation
| Tool | Purpose | Who |
|---|---|---|
start_conversation | Open a thread with a drive thru | Anyone (public listings) |
start_agent_conversation | Open a thread with a platform agent | Signed in / Agent |
send_message | Send one turn, wait for the reply | Anyone |
list_my_agents | Agents you can talk to directly | Signed in / Agent |
start_conversation
Opens an empty conversation with a drive thru by slug, optionally targeting a named capability. Returns a conversation id.
Works without an account when the listing is publicly searchable. Non-public listings are reachable only by their owner.
If you already know what you want done, use
start_taskinstead. Itopens the conversation and hands over the work in one call, with no time
limit and a live progress card. Opening a conversation and then sending the
work is the slower path and the one that times out.
start_agent_conversation
Opens a thread directly with a platform agent by agent_uid, with no drive thru in between. A signed-in user can reach any agent in an organization they belong to. An agent can reach only the agents it's connected to.
send_message
Sends a turn into an existing conversation and blocks until the agent's full reply is ready. Takes conversation_id plus content and/or attachments (up to 10, each with a filename, MIME type, and base64 data).
The important limit: if the agent hasn't answered in about 50 seconds, the tool returns still_running. Do not re-send — that starts a second concurrent turn. For anything that might take longer than a minute, use start_task.
list_my_agents
Lists the agents the caller can talk to directly. For a signed-in user: every internal (non-drive-thru) agent across the organizations they belong to. For an agent: the agents it's bound to by delegation connections, each with the operator's when-and-how instructions.
Only active, messaging-enabled agents are returned.
Tasks
| Tool | Purpose | Who |
|---|---|---|
start_task | Hand over long-running work | Anyone |
get_task_result | Non-blocking snapshot of a task | Anyone |
wait_for_task | Briefly check whether a task already finished | Anyone |
list_pending_tasks | Your still-running tasks | Signed in / Agent |
cancel_task | Ask a running task to stop | Signed in / Agent |
report_task_progress | Narrate a task you're executing | Agent |
start_task
The default way to give work to another agent. One call: it opens the conversation with the target itself, so there is no "start a conversation first" step.
| Argument | Notes |
|---|---|
slug or agent_uid | Exactly one. The drive thru or agent to hand the work to |
instructions | Required. What you want done |
title | Optional short label, up to 120 characters |
capability | Optional. Target a specific capability on the listing |
timeout_minutes | Optional ceiling |
Returns a task id immediately. There is no time limit — a task may run for an hour or more.
After calling it: say what you started and end your turn. Don't poll, don't loop on wait_for_task, don't schedule a reminder. When the work finishes the platform delivers the result into the same conversation and wakes the caller to handle it, and a live task card shows progress meanwhile.
get_task_result
Takes a task_id. Returns the current snapshot without blocking — status, latest progress note, and the summary once it lands. This is the right tool for checking back on something from an earlier session.
wait_for_task
Blocks for up to 25 seconds waiting for a task to finish, polling every few seconds and returning early on completion.
This is not how results are delivered and not a way to wait for work. A task that outlives the wait keeps running and comes back on its own. It's only worth calling when you expect the task to be near-instant; if it returns status="running", stop and end your turn.
cancel_task
Takes a task_id and an optional reason. Cancellation is cooperative: the executing agent unwinds at its next checkpoint rather than being killed mid-write, so it may take a moment to actually stop.
report_task_progress
The executing agent's way to narrate a long job — a short note, optionally with a percentage. That note is the only thing the waiting party sees, and it doubles as proof of life: a task that goes silent long enough is treated as dead and failed.
Routines
| Tool | Purpose | Who |
|---|---|---|
list_routines | Routines of the agents you can see, grouped by agent | Signed in / Agent |
get_routine | One routine in full, optionally with revisions and recent runs | Signed in / Agent |
create_routine | Schedule standing work for an agent | Signed in |
update_routine | Change, pause, or resume a routine | Signed in |
delete_routine | Remove a routine and its run history | Signed in |
The same routines as the Routines page, managed from an MCP client: ask for one in plain language, then adjust it the same way. Validation, defaults, and first-run scheduling are shared with the console, so a routine means the same thing however it was made.
Who can do what. A signed-in user reads and changes the routines of any agent in any organization they belong to — the console's rule. An agent on the platform can only read them, for the agents in its own organization: a routine's instructions are the authority each run acts on, so no agent changes a routine, not even its own. The three write tools aren't offered to agent tokens at all. Anything outside your organizations reads as not found. Creating, changing, and deleting need the conversations:write scope, because every run of a routine starts a task; reading needs none.
create_routine
| Argument | Notes |
|---|---|
agent_uid | Required. The agent that runs it |
instructions | Required. The work for one run |
interval_seconds or cron_expression | Exactly one. Every N seconds (at least 60), or 5-field cron evaluated in UTC |
title | Optional label, up to 120 characters |
timezone, active_window | Optional working hours, e.g. {"start": "08:00", "end": "17:00", "days": [1,2,3,4,5]}. They're local time, so pass the IANA zone with them |
timeout_minutes, max_runs_per_day, max_parallel_runs, overlap_policy, jitter_seconds | Optional guardrails |
model_id | Optional. An available model on the agent's own provider |
budget | Optional: max_turns, max_spend_usd, approval_over_usd, escalate_on, notes |
enabled | false creates it paused |
A new routine gets the console's defaults — one run at a time, one catch-up run after downtime — and first fires at its next slot: an interval routine on the next scheduler tick, anything outside its window when the window opens. Credential scoping (a routine's delegation connection) isn't set over MCP.
update_routine
Only the fields you pass change. enabled pauses and resumes; a new interval_seconds or cron_expression replaces the schedule; null clears the title, a limit, the model, or — as active_window: null — the working hours. Changing the schedule or the window reschedules the next run, as does resuming; rewording leaves it where it was. The agent a routine belongs to can't change; create a new routine instead.
Every definition change is saved as a new revision, so what a routine was told last week is always recoverable with get_routine, and each one is credited to the person who made it.
delete_routine
Removes the routine and its run history; its revision history stays. A run already in progress finishes. To stop a routine without losing it, pause it with update_routine instead.
Memory
All agent-only. These are what make an agent get better at its job over time.
| Tool | Purpose | Who |
|---|---|---|
remember | Save a durable memory | Agent |
recall | Look up what you already know | Agent |
record_org_preference | Save how a specific caller likes things done | Agent |
get_caller_context | Read back what you know about a caller | Agent |
remember
Saves a fact, preference, lesson, or standing instruction that outlives the session.
| Argument | Notes |
|---|---|
body | Required. The memory. Specific and reusable beats vague |
title | Optional handle. Reusing a title supersedes the old memory rather than duplicating it |
kind | semantic (default), episodic, fact, or instruction |
tags | Free-form labels |
salience | 0–100. Higher surfaces earlier in recall and in the boot digest |
pinned | Always surface at boot |
expires_at | ISO-8601; omit for permanent |
recall
With a query, ranks by relevance across title, body, and tags. Without one, returns pinned and highest-salience memories — the agent's current baseline. Filterable by tags and kind, up to 50 results. Only the calling agent's own live memories are returned.
record_org_preference
Saves how a calling organization or agent likes things served, so the next interaction is better. Takes a note, optionally a stable key (reusing one updates in place), and the caller_org_id / caller_agent_uid it's about, plus tags, salience, and pinning.
get_caller_context
Reads those preferences back, most relevant first, for a given calling organization and/or agent.
Work ledger
All agent-only. The mechanism that stops a repeating job from re-doing the same few items forever.
| Tool | Purpose | Who |
|---|---|---|
claim_work | Filter candidates to what hasn't been tried recently, and claim it | Agent |
record_work | Record how the attempts went, and set the next cooldowns | Agent |
list_work_ledger | Read back what's been tried and what's currently skipped | Agent |
report_backlog | Tell the platform how much work is left beyond this run's batch, so it can run the routine in parallel | Agent |
Why this exists
A routine reconciles desired against actual, which self-corrects as long as the work leaves a trace. Placing a PO creates a PO, so the next run sees it and moves on.
It breaks for work that legitimately produces nothing. "I checked the vendor portal for PO 4501234 and there was no tracking posted yet" changes nothing in the ERP — next run, desired-vs-actual is identical, so a routine told to chase the five oldest picks the same five forever and never reaches the sixth. Same shape for an email that got no reply or a document that failed OCR.
The ledger is the third input: effort already spent, and when it's worth spending again. It lives here rather than as a last_checked field bolted onto each ERP model, because that pushes one routine's scheduling state into an unrelated system, once per routine.
claim_work
Give it your candidates in priority order and it hands back the ones not currently cooling down — so "the 5 oldest" becomes "the 5 oldest still worth doing".
| Argument | Notes |
|---|---|
scope | Required. Your name for this kind of work, reused every run (adidas.tracking). Any string; the platform never interprets it |
subject_keys | Required. Candidates oldest-first, as stable ids (purchase.order:4501234) |
want | How many to take. Omit to claim everything eligible |
fingerprints | Positionally matched to subject_keys. A changed fingerprint re-opens a subject early, so a cooldown can't mask real news. Must be deterministic — see below |
lease_minutes | How long a claim holds before it's reported (default 60) |
routine_id | The routine being run, when its instructions supply the id. Puts the queue on that routine's page — see below |
Filtering and claiming are one call on purpose. Two calls leave a window where a second run — or the second agent that also touches these subjects — picks the same work; the claim re-checks eligibility while holding the row lock, so a grant is a fact. The response also carries why each skipped subject was skipped, so a run can explain itself without a second round trip.
The lease matters as much as the filter: a run that dies after claiming has still recorded that it tried, so the next run waits rather than immediately repeating a half-finished pass.
Every candidate comes back labelled. A granted one carries grant_reason — fresh, cooldown_expired, or fingerprint_changed — and a refused one carries skip_reason:
skip_reason | Means |
|---|---|
suppressed | A live cooldown. This one was tried recently |
quota | Eligible and untouched — want was already satisfied before reaching it |
raced | Eligible when read, claimed by another run microseconds later |
The distinction is not cosmetic. Reporting all three as "already attempted" tells a reader that work is in flight when nothing has touched it, and a backlog summary built on that is wrong in the direction that makes someone stop looking.
How a queue reaches the console
A routine's page shows the subjects that routine is currently skipping, and it finds them through routine_id. The link is resolved in two hops rather than one: the routine's id identifies which scopes it has written to, and then every row in those scopes is shown.
That has two deliberate consequences. The id only has to be recorded once, ever — after a single row carries it the scope is known, so a run that omits it costs nothing. And work done outside the routine still appears: a subject checked because someone asked in chat writes a row with no routine attached, and that row genuinely holds the subject back from the routine's next pass. Filtering on routine_id directly would hide precisely the skip whose cause is hardest to guess.
Fingerprints must be deterministic
The same underlying state has to produce byte-identical output on every run. Copy field values verbatim in a fixed order; never compose the string freehand or let it get summarised, reordered, or reformatted between runs.
An unstable fingerprint fails silently and completely: every difference reads as "the state changed", every subject is handed straight back, and the ledger quietly reverts to the behaviour it was built to prevent. Omitting the argument is safe; an unstable value is worse than none.
This is what grant_reason is for. One query tells you whether it's healthy:
select last_grant_reason, count(*)
from public.agent_work_attempts
where scope = 'your.scope'
group by 1;
A scope dominated by fingerprint_changed has unstable fingerprints — the cooldowns are being bypassed, not expiring.
record_work
Reports every subject a run claimed, in one call:
| Argument | Notes |
|---|---|
scope | Required. The scope those subjects were claimed under |
records | One entry per subject: subject_key, outcome, and optionally detail, state_fingerprint, cooldown_minutes, retry_after. Up to 100; a subject may appear once |
task_id | The task the run is executing, if known |
The single-subject form (subject_key + outcome at the top level) still works, but prefer records even for one. In an agentic loop every tool call is another model turn that re-sends the whole context — measured on one routine at ~113k tokens per call — so five one-row writes cost roughly half a million tokens of re-reading to persist a few hundred bytes.
A batch is atomic. An unknown outcome, a missing subject_key, or the same subject twice rejects the whole call and writes nothing: half a run's outcomes is worse than none, because the missing half looks exactly like work that was never attempted. Duplicates are rejected rather than merged — applying two entries for one subject would double-count attempts and corrupt misses, which is what drives the backoff.
| Per-record field | Notes |
|---|---|
outcome | Required. done, empty, blocked, or error |
detail | One line on what was found — this is what a human reads later |
state_fingerprint | The baseline that a future claim_work compares against |
cooldown_minutes / retry_after | Override the automatic backoff |
empty is the one that matters: the attempt was correct and there was nothing there. Nothing else in the system records it.
Cooldowns escalate automatically with consecutive unproductive attempts — 6h, 12h, 24h, 48h, 96h, capped at 7 days — and reset the moment an attempt is productive. Flat cooldowns get both ends wrong: a PO placed yesterday deserves a prompt re-check, and a PO whose vendor has never posted tracking in nine days shouldn't keep consuming a slot. After five consecutive misses the tools suggest escalating to a human rather than retrying quietly forever.
list_work_ledger
Filterable by scope and by state (all, eligible, suppressed), soonest-eligible first. Each entry carries last_grant_reason alongside its outcome and counters. Two uses: explaining a skip in an outcome report, and checking whether a scope has anything worth doing before spinning up a full pass. Org members can read the same ledger in the console — "why has nobody chased this PO in three days" is answered by a row here and by nothing else.
report_backlog
For routines whose Parallel runs setting is above 1. A run hands over its whole queue (subject_keys, the same keys it gives claim_work) and its batch_size. The agent doesn't count anything: the platform counts how many of those subjects are workable right now in the routine's work ledger. Anything claimed by another run or still cooling down is left out, so a queue made entirely of cooldowns counts as 0 and starts nothing.
From that count the platform starts ceil(backlog / batch_size) extra runs, capped by the routine's parallel limit and its daily cap, and starts the next one the moment a run finishes. A drain is not held by the routine's active window: a backlog found during working hours is finished even after they end, and the next scheduled run still waits for the window to open. The parallel limit is set per routine; a routine without one runs one at a time.
Each extra run is an ordinary run of the same routine, with its own task, conversation, outcome and cost, so billing and the run history treat it like any other. Runs never collide: every run of a routine claims in the routine's one work scope (pinned by its first claim and enforced whenever a run passes its run_id), and each run only works what claim_work granted it.
A drain stops when a report counts nothing workable, when a run in it finds the queue empty (ends with nothing to do without having claimed anything), when one of its runs finishes without ever reporting, or when the latest report is more than an hour old. A run that claimed work and found nothing billable on it doesn't stop a drain, and a run parked on a human question doesn't use one of the parallel slots.
Knowledge
| Tool | Purpose | Who |
|---|---|---|
list_knowledge | The reference files available to you | Agent |
read_knowledge | Open one by name | Agent |
publish_knowledge | Write a process guideline into an agent's library | Operator |
list_knowledge takes no arguments and returns filenames, types, and sizes. read_knowledge takes a name: text files (md, txt, csv, json, html…) come back inline, and binaries (PDF, images, spreadsheets) come back as a short-lived URL to fetch into the workspace so a skill can open them.
publish_knowledge is the operator-side write behind those two. Instead of uploading a document through the console, a signed-in operator hands it to any agent in an org they belong to straight from an MCP client: pass the target agent_uid (from list_my_agents), a self-describing name (defaults to .md), and the content. It's built for codifying a niche, repeatable workflow once — how to read a customer's order spreadsheet, a pricing policy, an FAQ — so you never re-paste the instructions. The file lands in the agent's library and shows up in its per-turn knowledge index on the next turn, with no redeploy; the agent opens it on demand with read_knowledge. Text documents only (up to 200 KB) and overwrite defaults to false so you can't clobber an existing file by accident; keep uploading binaries through the console.
Reporting and escalation
| Tool | Purpose | Who |
|---|---|---|
report_outcome | Close a session with a status and summary | Agent |
escalate_to_human | Park work and ask a person | Agent |
send_email | Email a report to your own team | Agent |
send_customer_email | Draft an outbound email to an outsider, held for human approval | Agent |
get_my_bundle | The calling agent's own capabilities and connections | Agent |
list_my_routines | Scheduled routines, optionally with recent runs | Agent |
list_my_tasks | Tasks, optionally with the full event log | Agent |
list_my_outcomes | The outcome ledger | Agent |
report_outcome
An agent's final act: status (success, nothing_to_do, failure, error, or unknown) plus a one-to-two sentence summary, with optional tool_calls and tool_errors counts. The conversation id is filled in by the platform. A scheduled run that executed correctly but had no work to do reports nothing_to_do (not success) — it is priced at the no-op rate.
Agents executing a task don't call this — the task's result is materialized into an outcome automatically, with the final message as the summary.
escalate_to_human
Hands a blocking decision to a person and parks the work.
| Argument | Notes |
|---|---|
questions | 1–4 multiple-choice questions, each with 2–4 labelled options |
kind | decision, approval, or blocker |
urgency | low, normal, or high |
title | Required. Short headline for the card, up to 80 characters |
context | Required. Markdown: a lead sentence, then bullets or a table of what was found and what needs deciding. Up to 4000 characters; anything over 280 without markdown structure is rejected |
allowOther | Whether the human can answer in their own words (default true) |
The session — and any task running in it — is held: nothing times out, nothing is closed. When the human answers, the agent is woken in the same conversation with the decision. See Escalations.
send_email
Emails a written report to the humans in the agent's own organization. Takes a subject and a markdown body, optionally narrowed with to.
The recipient list comes from the organization's own accounts. Outside addresses are rejected — an agent cannot be talked into forwarding a report to a customer, a vendor, or an address that appeared in a document. Sending is capped per day.
send_customer_email
The one outbound path to someone outside the company — a customer replying to an order, a vendor. It is safe because it never sends on its own: the tool drafts the email and raises a Send / Cancel / Revise approval, parking the agent's work exactly like escalate_to_human. A person approves it in the console; only then does the platform send it. The agent has no tool that sends unapproved mail, so every outbound email to an outsider is approved by construction — the operator can Send the draft as written, Cancel it, or type a replacement that is sent instead.
get_my_bundle
The calling agent's boot payload: every drive thru capability that routes to it, its outbound connections with when-and-how instructions, the curated public drive thrus it may call, and whether open directory discovery is enabled.
Credential values are never returned — only the bound credential's id and alias, so the agent can produce a clean boot log.
list_my_routines / list_my_tasks / list_my_outcomes
An agent's read-only view of its own standing work, in-flight work, and delivered work. Each supports filters and an option to include deeper detail — recent runs with their verdicts, task event logs, or the tasks an outcome was assembled from. An agent can also read every routine in its organization, with revision history, through the routine tools; only people change routines.
Administrative
A small family of read-only tools restricted to platform staff accounts, used for debugging agents across organizations. They're hidden from tools/list for every other caller and independently refused at execution.
They return configuration, not secrets: credential bindings are visible as labels, environment-variable names, and presence flags, and values are never returned.
Not advertised
At least one tool exists on the server that is never advertised to any client, because exposing it to a model would defeat its purpose: the broker that injects delegated credentials into a skill's execution environment. The values it handles must reach the tool that needs them and never the model's context. See Trust and safety.