{"title":"Permission to act.","description":"Inspect what an agent asks to do, what the runner allows, and what actually executes. A small offline harness with explicit boundaries.","theme":"teal","kind":"Agentic AI","cases":[{"id":"normal","name":"A valid offline task","title":"A tool request is not permission.","unit":"Cumulative event counts","parameter":"Decision step","x":[1,2,3],"series":[{"label":"Validated executions attempted","values":[1,2,2]},{"label":"Blocked / failed requests","values":[0,0,0]}],"snapshots":[{"rows":[[1,"lookup_user","executed","{'name': 'Linh', 'tier': 'gold'}"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"lookup_user","executed","{'name': 'Linh', 'tier': 'gold'}"],[2,"calc","executed","{'result': 50.0}"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"lookup_user","executed","{'name': 'Linh', 'tier': 'gold'}"],[2,"calc","executed","{'result': 50.0}"],[3,"final","answer","The fictional user is Linh; the arithmetic result is 50."]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."}],"columns":["Step","Requested tool","Result","Observation"],"context":"Offline MockLLM, six decision steps maximum, three validated tool attempts. Only calculator and fictional-user lookup are granted by default.","readout":"Inspect the trace, not only the final answer. The draft tool can only change an in-memory list and is denied in this showcase. No network, shell, real data or credentials are used.","note":""},{"id":"permissions","name":"A script asks for more authority","title":"A tool request is not permission.","unit":"Cumulative event counts","parameter":"Decision step","x":[1,2,3,4],"series":[{"label":"Validated executions attempted","values":[0,0,0,0]},{"label":"Blocked / failed requests","values":[1,2,3,3]}],"snapshots":[{"rows":[[1,"save_draft","blocked","PermissionError: Write capability not granted"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"save_draft","blocked","PermissionError: Write capability not granted"],[2,"shell","blocked","ValueError: Unknown tool"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"save_draft","blocked","PermissionError: Write capability not granted"],[2,"shell","blocked","ValueError: Unknown tool"],[3,"lookup_user","blocked","ValueError: Exact argument names required"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"save_draft","blocked","PermissionError: Write capability not granted"],[2,"shell","blocked","ValueError: Unknown tool"],[3,"lookup_user","blocked","ValueError: Exact argument names required"],[4,"final","answer","Blocked calls did not execute. No write or shell capability was granted."]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."}],"columns":["Step","Requested tool","Result","Observation"],"context":"Offline MockLLM, six decision steps maximum, three validated tool attempts. Only calculator and fictional-user lookup are granted by default.","readout":"Inspect the trace, not only the final answer. The draft tool can only change an in-memory list and is denied in this showcase. No network, shell, real data or credentials are used.","note":""},{"id":"budget","name":"A loop reaches its budget","title":"A tool request is not permission.","unit":"Cumulative event counts","parameter":"Decision step","x":[1,2,3,4,5,6],"series":[{"label":"Validated executions attempted","values":[1,2,3,3,3,3]},{"label":"Blocked / failed requests","values":[0,0,0,1,2,3]}],"snapshots":[{"rows":[[1,"calc","executed","{'result': 2}"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","executed","{'result': 2}"],[2,"calc","executed","{'result': 2}"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","executed","{'result': 2}"],[2,"calc","executed","{'result': 2}"],[3,"calc","executed","{'result': 2}"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","executed","{'result': 2}"],[2,"calc","executed","{'result': 2}"],[3,"calc","executed","{'result': 2}"],[4,"calc","blocked","PermissionError: Tool-call budget exhausted"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","executed","{'result': 2}"],[2,"calc","executed","{'result': 2}"],[3,"calc","executed","{'result': 2}"],[4,"calc","blocked","PermissionError: Tool-call budget exhausted"],[5,"calc","blocked","PermissionError: Tool-call budget exhausted"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","executed","{'result': 2}"],[2,"calc","executed","{'result': 2}"],[3,"calc","executed","{'result': 2}"],[4,"calc","blocked","PermissionError: Tool-call budget exhausted"],[5,"calc","blocked","PermissionError: Tool-call budget exhausted"],[6,"calc","blocked","PermissionError: Tool-call budget exhausted"]],"note":"Final runner state: step_limit. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."}],"columns":["Step","Requested tool","Result","Observation"],"context":"Offline MockLLM, six decision steps maximum, three validated tool attempts. Only calculator and fictional-user lookup are granted by default.","readout":"Inspect the trace, not only the final answer. The draft tool can only change an in-memory list and is denied in this showcase. No network, shell, real data or credentials are used.","note":""},{"id":"invalid","name":"Malformed or expensive arithmetic","title":"A tool request is not permission.","unit":"Cumulative event counts","parameter":"Decision step","x":[1,2,3,4],"series":[{"label":"Validated executions attempted","values":[1,2,3,3]},{"label":"Blocked / failed requests","values":[1,2,2,2]}],"snapshots":[{"rows":[[1,"calc","blocked","ValueError: Only finite arithmetic with + - * / is supported"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","blocked","ValueError: Only finite arithmetic with + - * / is supported"],[2,"calc","blocked","ZeroDivisionError: division by zero"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","blocked","ValueError: Only finite arithmetic with + - * / is supported"],[2,"calc","blocked","ZeroDivisionError: division by zero"],[3,"calc","executed","{'result': 21}"]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."},{"rows":[[1,"calc","blocked","ValueError: Only finite arithmetic with + - * / is supported"],[2,"calc","blocked","ZeroDivisionError: division by zero"],[3,"calc","executed","{'result': 21}"],[4,"final","answer","Only the bounded expression completed."]],"note":"Final runner state: answered. Tool budget counts attempted validated executions, including runtime failures. Blocked requests do not call a tool. This script is deterministic; no real LLM was evaluated."}],"columns":["Step","Requested tool","Result","Observation"],"context":"Offline MockLLM, six decision steps maximum, three validated tool attempts. Only calculator and fictional-user lookup are granted by default.","readout":"Inspect the trace, not only the final answer. The draft tool can only change an in-memory list and is denied in this showcase. No network, shell, real data or credentials are used.","note":""}],"config":{"maxSteps":6,"maxCalls":3,"realLLM":false,"writeGranted":false},"method":["Reuses the curriculum's MockLLM and Tool interface, with a separate bounded runner. Allowed tool names, exact argument names and short string types are checked before execution.","The calculator parses an AST instead of eval: only +, −, ×, ÷ and finite bounded values. Input length and node count are limited. Validated execution attempts consume the tool budget even if the operation fails.","Scripts cover normal execution, missing authority, malformed requests and repeated calls. The downloadable JSON contains the step-by-step observations."],"limits":["This does not measure a real model's intelligence, prompt-injection resistance or factuality. Scripted negative tests verify these particular code paths only.","This is not a general JSON Schema validator, process sandbox, persistent agent, or production authorization system. A real write workflow needs explicit user authorization and durable audit/confirmation."],"references":[],"provenance":{"python":"3.13.0","numpy":"2.4.6","sourceSha256":{"showcase_benchmark.py":"69d3b0dbdffaba760ab41baf0a932c29194337f14b7e488753b27ebe32081078","offline_guard.py":"f1cc1f347e061680948809a3d8a26c6aa59acaa58a62f043a016f34cc7c80881","agent.py":"3986524405cc7b4368e921afb80afc7abe3523c9621510989a920f06a260796b"}}}
