Try:
Tools available to the agent: calculator (safe arithmetic) and run_python (a restricted Python sandbox — AST-validated, no file or network access).
Pick an example or write your own task. The agent runs a plan → act → observe loop and shows its full reasoning trace.