Skip to main content

Command Palette

Search for a command to run...

Your agent shouldn't know about tools it can't use

Updated
•16 min read•View as Markdown
Your agent shouldn't know about tools it can't use
J
AWS cloud architect and AI enthusiast based in Morgantown, WV. I write about building on AWS, exploring AI agents and automation, and the occasional deep dive into cloud security. Follow along as I share what I'm learning, building, and breaking in the cloud

I gave a clinical operations user access to two tools on an Amazon Bedrock AgentCore Gateway. Semantic tool search still returned all six tools, with full schemas, on every query. The gateway's Cedar policies filter tools/list per caller, but search results aren't filtered by policy. So treat search as a relevance ranker, not an authorization layer, and intersect the two in your agent before any tool reaches the model.

This post walks you through the pattern end to end with a Strands agent, real healthcare APIs and Cedar policies in ENFORCE mode, plus what broke along the way.

The problem: one agent, every tool

Most Strands tutorials hand the agent a static list of tools and an AWS Identity and Access Management (IAM) role broad enough to call all of them. That works with five tools. It stops working at two hundred, for two reasons.

The first is cost. Every tool schema rides along on every request. In my measurements, each schema cost about 174 input tokens on Amazon Nova Pro, so a 1,000-tool catalog adds roughly 174,000 tokens to every request before the user types a word.

The second is the audit. Life sciences companies work under GxP, the good practice regulations such as GMP and GCP. In that environment, someone will ask which tools the agent can reach, on whose behalf, and who decided. "The prompt tells it not to" isn't an answer a GxP auditor accepts. The boundary has to sit outside the model, behave deterministically and leave a log.

The alternative is to stop giving the agent a catalog. Instead, the agent asks the gateway two questions per request: what am I allowed to use, and what is relevant here?

The architecture: discover per request, enforce per call

The gateway exposes tools over the Model Context Protocol (MCP), and the agent never holds a catalog. For each request, the agent runs four steps, and the gateway's Cedar policy engine makes the final decision on every call.

Per-request tool discovery through AgentCore Gateway
  1. tools/list: what may this caller see? The gateway evaluates the caller's JSON Web Token (JWT) against the policy engine (PartiallyAuthorizeActions, with no arguments) and returns only the tools some policy could permit.

  2. Semantic search: what is relevant? The built-in x_amz_bedrock_agentcore_search tool ranks tools by similarity to the question.

  3. Intersect. Keep tools that are both relevant and listed, top four by search rank. If the top search result is a tool the caller can't see, tell the model to decline.

  4. Run the model with only those schemas. Every tools/call is still checked by the policy engine (AuthorizeAction), this time with the real arguments.

Step 3 is the part I didn't expect to need. I assumed search would respect the same policies as tools/list, because both go through the same gateway with the same caller token. It doesn't. Without step 3, the unfiltered search undoes the filtered list.

Real tools, three personas

Mock tools would hide the problems, so the demo uses six real ones behind three AWS Lambda targets. Two targets call public FDA and NIH APIs. The third writes permanent records to an Amazon DynamoDB table.

Target Tool Backed by Writes?
openfda search_adverse_events FDA Adverse Event Reporting System (FAERS) No
openfda search_recalls FDA recall enforcement reports No
openfda get_drug_label FDA structured product labeling No
ctgov search_trials ClinicalTrials.gov API v2 No
pv list_safety_signals DynamoDB signal log No
pv log_safety_signal DynamoDB signal log Yes

Three Amazon Cognito users stand in for three roles. A pre-token generation Lambda function flattens each user's Cognito group into a single string persona claim, which Cedar reads as principal.getTag("persona").

Persona Allowed Notable restriction
clinical_ops search_trials, get_drug_label No adverse events, recalls or signal log
safety_analyst All six tools Cannot log a critical severity signal
safety_lead All six tools None; the only persona that can log critical

The analyst's restriction is the interesting one. It depends on an argument, so the gateway can't account for it when it builds the tool list. The next sections show how to wire these rules into the gateway, then what happened when I tested them.

Prerequisites

To follow along, you need:

  • An AWS account and a Region where Policy in AgentCore is available. I used us-east-1.

  • Model access in Amazon Bedrock for the model you set as MODEL_ID. The repo defaults to Amazon Nova Pro.

  • Python 3.13 or later, uv, and Node.js for the AWS CDK CLI.

  • An AWS CDK bootstrapped account (npx cdk bootstrap).

  • IAM permissions to create AgentCore gateways, targets, policy engines and policies, including bedrock-agentcore:InvokeGateway (explained below).

Everything in the demo is pay-per-use: Lambda, DynamoDB on-demand, Cognito, AgentCore Gateway and Policy, and Bedrock tokens. Costs stay small for a few test runs, but follow the clean-up steps at the end when you're done.

Setting it up: order matters

The AWS Cloud Development Kit (AWS CDK) stack deploys the Lambda functions, DynamoDB table, Cognito user pool and gateway execution role. You then create the AgentCore resources with scripts/setup_gateway.py, in this order, because each step depends on the one before:

  1. Create the policy engine. The gateway needs its ARN at creation.

  2. Create the gateway with semantic search on and policy in LOG_ONLY mode. You can only enable semantic search when you create the gateway. You can't add it later.

  3. Add the targets. The policy engine generates its Cedar schema from the targets' tool definitions.

  4. Create the Cedar policies, permits first. Policies are validated against that schema, and a forbid is rejected until a matching permit exists.

  5. Switch to ENFORCE mode after the LOG_ONLY decision logs look right.

The gateway call carries both the search switch and the policy engine:

ctl.create_gateway(
    name="discovery-demo-gateway",
    roleArn=outputs["GatewayRoleArn"],
    protocolType="MCP",
    authorizerType="CUSTOM_JWT",
    authorizerConfiguration={"customJWTAuthorizer": {
        "discoveryUrl": outputs["DiscoveryUrl"],
        "allowedClients": [outputs["AppClientId"]],
    }},
    protocolConfiguration={"mcp": {"searchType": "SEMANTIC"}},  # creation time only
    policyEngineConfiguration={"arn": engine_arn, "mode": "LOG_ONLY"},
)

A permit scopes each persona to its tools. Action names are <target>___<tool>:

permit (
  principal is AgentCore::OAuthUser,
  action in [
    AgentCore::Action::"ctgov___search_trials",
    AgentCore::Action::"openfda___get_drug_label"
  ],
  resource == AgentCore::Gateway::"<gateway-arn>"
)
when {
  principal.hasTag("persona") &&
  principal.getTag("persona") == "clinical_ops"
};

The forbid reads the tool's input. Forbid wins over permit, so it overrides the analyst's write permission for one severity only:

forbid (
  principal is AgentCore::OAuthUser,
  action == AgentCore::Action::"pv___log_safety_signal",
  resource == AgentCore::Gateway::"<gateway-arn>"
)
when { context.input has severity && context.input.severity == "critical" }
unless { principal.hasTag("persona") && principal.getTag("persona") == "safety_lead" };

Check two IAM details before you start, because missing either one can cost you an afternoon. The gateway execution role needs bedrock-agentcore:AuthorizeAction, PartiallyAuthorizeActions and GetPolicyEngine, or every call is denied. The role running setup needs bedrock-agentcore:InvokeGateway, because CreatePolicy validates each action by calling the gateway.

What I measured: the list is filtered, search results aren't

I ran a discovery matrix against the live gateway in us-east-1 with the policy engine in ENFORCE: three personas, four questions each, plus five call-time probes.

tools/list is filtered per caller. clinical_ops saw exactly its two tools. Both safety personas saw all six. In LOG_ONLY every persona saw everything, so a LOG_ONLY run tells you nothing about filtering.

Search results aren't filtered by policy. For clinical_ops, search returned all six tools on every query, ranked by relevance. The ranking was identical for all three personas.

Query (as clinical_ops) Top search result Hidden by tools/list, returned by search
how many serious adverse event reports are there for semaglutide search_adverse_events 4 tools
find recruiting phase 3 trials for type 2 diabetes search_trials 4 tools
has metformin been recalled search_recalls 4 tools
record a new safety signal for review log_safety_signal 4 tools

That's 16 hidden tools surfaced in four queries. Search returns each tool's name, description and input schema, so the agent can't call them, but the model now sees that they exist and how to call them. For search_adverse_events, a tool clinical_ops can't use, that's the full definition as registered on the gateway:

{
  "name": "openfda___search_adverse_events",
  "description": "Summarize post-marketing adverse event reports (FDA FAERS) for a drug: total report count and the most frequently reported reactions (MedDRA preferred terms). Use for pharmacovigilance and safety signal questions about a marketed drug.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "drug_name": {"type": "string", "description": "Generic or brand name, e.g. 'semaglutide' or 'Ozempic'."},
      "serious_only": {"type": "boolean", "description": "Only count reports flagged serious (death, hospitalization, disability, life-threatening)."},
      "top_n": {"type": "integer", "description": "Number of top reactions to return, 1-25. Default 10."}
    },
    "required": ["drug_name"]
  }
}

The same run turned up two more findings:

  • The search tool needs no permit. Any authenticated caller can use it. Under ENFORCE it drops out of tools/list but stays callable.

  • Search ranks; it does not cut. With six tools it returned the whole catalog every time. The agent applies its own top-k.

Listed does not mean authorized. The call-time probes show the second layer:

Persona Call Result
clinical_ops search_adverse_events Denied: no policy applies (default deny)
clinical_ops search_trials Allowed
safety_analyst log_safety_signal, severity critical Denied by critical_signals_lead_only
safety_analyst log_safety_signal, severity low Allowed
safety_lead log_safety_signal, severity critical Allowed

The analyst's tools/list includes log_safety_signal, because listing is evaluated without arguments: a tool is listed if any call to it could be permitted. The forbid only bites when the real arguments arrive. The denial comes back as a JSON-RPC error that names the deciding policy:

{"jsonrpc": "2.0", "id": 1, "error": {"code": -32002, "message":
 "Tool Execution Denied: Tool call not allowed due to policy enforcement [Policy evaluation denied due to critical_signals_lead_only-bk9zn__6ww]"}}

The agent: intersect, then tell the model what's missing

Because search won't filter for you, the agent has to. The Strands agent connects to the gateway with the caller's token, then builds its tool set for each request. The intersection takes a few lines:

visible = {t.tool_name: t for t in all_visible_tools(client) if t.tool_name != SEARCH_TOOL}
relevant = search(client, question)  # ranked names from x_amz_bedrock_agentcore_search
selected = [visible[name] for name in relevant if name in visible][:MAX_TOOLS]
best_match_withheld = bool(relevant) and relevant[0] not in visible

If the search call itself fails, the agent falls back to everything tools/list returned. That list is already policy-filtered, so the fallback only costs the token saving.

I ran it on Amazon Nova Pro, mostly because Anthropic models weren't enabled in my account yet. That turned out to help: if a smaller, lower-cost model stays inside the boundary, the boundary isn't depending on the model's judgment.

Persona Question Outcome
clinical_ops Any recruiting phase 3 semaglutide trials in obesity? Answered from ClinicalTrials.gov
clinical_ops What are the top serious adverse events for semaglutide? "I don't have access to the data needed… adverse event report data"
clinical_ops Has metformin been recalled recently? "I don't have access to recall data"
safety_analyst What are the top serious adverse events for semaglutide? Answered from FAERS: nausea 4,911 reports, vomiting 4,285
safety_analyst Log a critical signal for drug X Tool called, gateway denied it, agent reported "not permitted"
safety_lead Log a critical signal for drug X Signal written to DynamoDB

Hiding a tool isn't enough on its own. With a generic "say so if you lack a tool" prompt, the clinical_ops agent answered the adverse event question from the warnings section of the drug label.

That's a plausible answer from the wrong source, which in pharmacovigilance is worse than no answer. A stricter prompt fixed it two times out of three.

What made it reliable, four runs out of four, was a deterministic signal. When the top search result is a tool the caller can't see, the agent adds one instruction: the best-matching capability is unavailable, so decline. It never names the hidden tool. With that instruction, the same question gets a clean refusal:

"I don't have access to the data needed… adverse event report data"

The model reaches for the nearest tool it has unless you tell it the right one exists and is out of reach.

What discovery saves

Input tokens per request, measured with Bedrock Converse on Nova Pro:

Query Tools sent Input tokens Full catalog Change
serious adverse event reports for semaglutide 4 of 6 1,005 1,410 −29%
recruiting phase 3 trials for type 2 diabetes 4 of 6 1,101 1,408 −22%
has metformin been recalled 4 of 6 1,041 1,401 −26%
record a new safety signal 4 of 6 994 1,404 −29%

Each tool schema cost about 174 tokens, plus a fixed 353 tokens whenever any tools are sent. Search added roughly half a second per request (438 to 516 ms).

With six tools, the saving is modest. The per-tool cost is what scales. The following table extrapolates from those measurements, assuming schemas of similar size. It's arithmetic, not a measurement.

Catalog size All tools in context Search-scoped (4 tools)
50 9,053 tokens 1,049 tokens
200 35,153 tokens 1,049 tokens
1,000 174,353 tokens 1,049 tokens

Gotchas from a real deployment

None of these showed up until the stack ran against a real account.

  • Permits before forbids. CreatePolicy rejected the forbid as "overly restrictive" until a permit existed for the same principal type and action. The setup script now creates permits first.

  • Scope the forbid's principal. A bare principal also covers AgentCore::IamEntity, which has no permit, so validation flags it. principal is AgentCore::OAuthUser fixes it.

  • Denials are JSON-RPC errors (code −32002), not the isError tool results the docs example shows. Handle both.

  • The gateway throttles bursts. A tight loop of search calls returned HTTP 429. Retry with backoff.

  • LOG_ONLY hides filtering. It logs what would be denied but lists every tool to everyone, so test discovery in ENFORCE.

  • Cedar sees claims as string tags. Cognito's cognito:groups claim is an array, so a pre-token trigger flattens it into one persona string. Customizing access tokens this way needs the Cognito Essentials plan.

  • openFDA widens silently. A parenthesized OR across two fields, such as (generic_name:"x"+brand_name:"x"), matched all 20.7 million reports instead of the drug's. The tools query one field at a time.

  • Anthropic models on Bedrock need a use-case form submitted for the account before the first call.

Clean up

To avoid ongoing charges, delete the AgentCore resources first, then the CDK stack:

uv run scripts/teardown.py   # policies, targets, gateway, policy engine, demo users
cd infra && npx cdk destroy   # Lambda functions, DynamoDB table, Cognito user pool, IAM role

The teardown script deletes in reverse order of creation, because a gateway can't be deleted while its targets exist.

Why this matters in regulated environments

With the rough edges covered, here's why the pattern is worth the effort for regulated teams. Most agent frameworks leave authorization to application code. Policy in AgentCore moves it outside the model, into a reviewable policy language, and logs every decision.

  • The rules are auditable artifacts. Four Cedar files say who can call what, under which conditions. A quality reviewer can read them without reading agent code, and they live in version control next to the infrastructure.

  • LOG_ONLY maps to a validation phase. Run the policies against real traffic, review the decision logs, then switch to ENFORCE. That mirrors a computer system validation (CSV) step.

  • Every decision is traceable. The signal log stores the AgentCore request and MCP message IDs on each record, so a written signal joins back to the gateway's policy decision, which holds the caller and the result.

If you take one thing into a design review, make it this: gateway policy filters tools/list and enforces every call, but semantic search results can describe tools the caller can't use. Treat search as a relevance ranker, not an authorization layer, and intersect in the agent.

Caveats and what's next

This was one Region, one day, six tools and four queries: enough to show the behavior, not to characterize it. The documentation doesn't say either way whether search results are policy-filtered, and that may change, so rerun the matrix before you rely on it. Agent outcomes are single runs on Nova Pro, apart from the decline case. FAERS counts are spontaneous reports and don't establish causation or incidence.

Next I'm putting a semantic layer behind the same gateway as MCP tools, and comparing a System 1 / System 2 routing design against an agent where the model makes every routing decision.

The code, Cedar policies, CDK stack and raw results are on GitHub: jrgwv/gateway-discovery-demo. Everything here reproduces with demo/discovery_matrix.py, demo/token_compare.py and agent/agent.py.

The views in this post are my own and don't necessarily represent those of my employer.

Sources