Fewer Tokens, Better Action:
GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

1Peking University 2National University of Singapore 3NVIDIA 4Impossible Research
†Corresponding author
Peking University National University of Singapore NVIDIA Impossible Research

The project in 80 seconds, with narration.

Success rate on nine sub-suites (radar), input tokens and cost per solved episode for tool calling and PyRUA-Lean

PyRUA-Lean matches or exceeds tool calling’s success rate on each sub-suite and costs less on every benchmark. (a) Success rate on all 700 instances, tagged by task property (semantic: perturbed layouts, objects, goals); RoboTwin 2.0 split by skill from task names. (b, c) Input tokens and GPT-6 Astra list-price dollars per solved episode, on instances both agents solved; Avg: mean over all of them. RC365: RoboCasa365.

Same planner, same primitives: code instead of tool calls

Success rate

71.7%vs 63.1% with tool calling

Input tokens

65% fewer276k vs 788k per solved episode

LLM calls

49% fewer8.7 vs 17.0 per solved episode

Cost

2.2× cheaper$0.74 vs $1.63 per solved episode

700 simulated task instances from LIBERO-PRO, RoboTwin 2.0 and RoboCasa365; tokens, calls and cost on the instances both agents solved.

Stop paying one LLM call per step

A robot agent built on a vision-language model usually acts through tool calling. The model picks one tool, such as move_to or release; the robot runs it; the result comes back, with three camera images after every motion; and the model is called again with everything so far. Every small step costs a full LLM call, and every call reads the growing conversation again, images included.

PyRUA-Lean (Python for Lean Robot-Use Agents) keeps the model, the robot primitives and the VLA policies, and changes only how the agent acts. The agent gets one tool, python(code), and a robot object robo whose methods are those same primitives. In one LLM call it writes a cell: a few lines of Python that chain several steps, check each result, retry when a policy falls short, and compute what no tool returns. The cell runs on the robot; the agent then sees only what the cell printed and the camera images it asked for.

Tool calling versus PyRUA-Lean, and a recorded placement example

PyRUA-Lean combines primitive composition with selective feedback. (a) The tool-calling agent invokes robot primitives through tool calls. Operations that depend on preceding results generally require another VLM turn, and motion calls automatically return images and state. (b) PyRUA-Lean instead generates Python cells that compose robot primitives with helper functions, conditional checks, and local retries. Intermediate execution stays within the runtime, while explicitly requested images and state messages are recorded and returned at cell end for replanning. (c) A recorded placement example illustrates both mechanisms: one PyRUA-Lean cell replaces baseline steps 14–17, reducing four LLM turns to one.

One cell instead of four LLM calls

This is the cell of panel (c), exactly as GPT-6 Astra wrote it while placing a bowl on a plate. Tool calling needed four LLM calls for the same step: two to locate the bowl and sample its base, one to lower it and one to release it, and each move returned three camera images. The cell measures the bowl's bottom in the point cloud, lowers it to the right height, releases it only if the move arrived, and asks for an image only if the task is still not done. The three comment lines are ours; plate was found by an earlier cell.

# compose: compute a placement target from the scene geometry
held_at_plate = robo.segment(point=(314, 795))
print('bowl above plate', held_at_plate)
assert held_at_plate.found
correction = np.array(plate.world_xyz[:2]) - np.array(held_at_plate.world_xyz[:2])
place_xy = np.array(robo.state().eef_pos[:2]) + correction
surface_map = robo.world_map()
bowl_patch = surface_map[290:365, 740:850]
bowl_heights = bowl_patch[:,:,2]
bowl_heights = bowl_heights[(bowl_heights>1.04) & (bowl_heights<1.17)]
bottom_z = float(np.quantile(bowl_heights, 0.02))
placement_z = plate.world_xyz[2] + robo.state().eef_pos[2] - bottom_z + 0.007
print('correction', correction, 'bottom', bottom_z, 'placement z', placement_z)
assert np.linalg.norm(correction)<0.07 and 0.95<placement_z<1.05
# compose: lower, check, and release within the cell
lower = robo.move_to([*place_xy, placement_z], gripper=+1, step_clip=0.012, tol=0.005, max_steps=100)
print('lower', lower)
if lower.reached and not robo.done:
    print('release', robo.release())
print('done', robo.done)
# select: ask for an image only if the task is unfinished
if not robo.done:
    robo.show('agentview')

It printed five lines and asked for no image: the bowl was on the plate and the task was done.

One task, two agents

The whole episode behind that example, run once with each agent on the same LIBERO-PRO task instance: “Pick the akita black bowl on the stove and place it on the plate”. Step through it one LLM call at a time: each square is one call, a filled square brought camera images back, and the outlined square solved the task.

What the agent writes

We read the code agent's cells on every benchmark, in solved and failed episodes, and counted what they do over all 11,145 cells of the 700 episodes. Four kinds of cells keep coming back, and none of them fits in one tool call. Each example is a real cell, exactly as the model wrote it.

Guarded chains

Several primitives in a row, each run only if the one before worked: move above the bowl, grasp with the VLA policy only if the move arrived and the task is not done, and look only if it is still not done. With tool calling, each link of the chain is an LLM call.

assert bowl.found and plate.found
bowl_xy = np.array(bowl.world_xyz[:2])
plate_xyz = np.array(plate.world_xyz)
table_z = robo.back_project(837,660).world_xyz[2]
prepose = [float(bowl_xy[0]),float(bowl_xy[1]),table_z+0.22]
approach = robo.move_to(prepose,gripper=robo.OPEN)
print('APPROACH',approach)
if approach.reached and not robo.done:
    picked = robo.pi0_pick('pick up the black bowl between the plate and the ramekin', max_chunks=20)
    print('PICK',picked)
print('STATE',robo.state())
if not robo.done:
    robo.show('agentview')
    robo.show('wrist')

Cell 2 · LIBERO-PRO, spatial swap, task 0, seed 0

Retries

Where a VLA policy may stop short of the goal, the cell runs it again until the task reports success. With tool calling, every retry is a round trip that brings back three camera images.

print(robo.task)
print(robo.success_criteria())
for attempt in range(3):
    result = robo.rldx_arm()
    print(result)
    print(robo.state())
    if robo.done or result.status != 'cap':
        break
if not robo.done:
    robo.show('agentview')
    robo.show('wrist')

Cell 1 · RoboCasa365 atomic, CloseToasterOvenDoor, seed 1

Perception of its own

The cell reads a camera's image and point cloud as arrays and locates objects itself: a colour mask for a block, the points above a height for its top, their principal axis for its orientation, and from that the yaw of the grasp. A tool-calling agent gets only what its tools compute.

mask=red&(rows<92)&(cols>98)&(cols<165)&(world[...,2]>0.745)
points=world[mask]
upper=points[points[:,2]>0.794]
mean_xy=np.mean(upper[:,:2],axis=0)
eigenvalues,eigenvectors=np.linalg.eigh(np.cov(upper[:,:2].T))
long_axis=eigenvectors[:,1]
if long_axis[1]<0: long_axis=-long_axis
short_axis=np.array([long_axis[1],-long_axis[0]])
projection_long=upper[:,:2]@long_axis
projection_short=upper[:,:2]@short_axis
block_center=long_axis*((np.min(projection_long)+np.max(projection_long))/2)+short_axis*((np.min(projection_short)+np.max(projection_short))/2)
print('flat centre',block_center,'long axis',long_axis,'extents',np.ptp(projection_long),np.ptp(projection_short))
left_grasp_xy=block_center+0.045*long_axis
right_grasp_xy=block_center-0.05*long_axis
yaw=math.atan2(long_axis[1],long_axis[0])
q_down=np.array([math.cos(yaw/2)/math.sqrt(2),-math.sin(yaw/2)/math.sqrt(2),math.cos(yaw/2)/math.sqrt(2),math.sin(yaw/2)/math.sqrt(2)])
print('grasp',left_grasp_xy,'receive',right_grasp_xy,'quat',q_down)
safe=np.array(robo.state().left.eef_pos);safe[2]=1.03
if checked_move('left',safe,quat=robo.state().left.eef_quat,substeps=20):
    waypoint=np.array([left_grasp_xy[0],0.015,1.03])
    if checked_move('left',waypoint,quat=q_down,substeps=25):
        print(robo.move_to('left',[left_grasp_xy[0],left_grasp_xy[1],0.97],quat=q_down,substeps=25))
print(robo.state().left)
robo.show('head')

Cell 5 · RoboTwin 2.0 without the VLA policy, handover block, seed 100003. checked_move is a helper the agent defined in an earlier cell.

Skills of its own

The agent writes a descent that moves down in 8 mm steps and stops as soon as a step does not reach its target, stops going down, or drifts sideways, and reuses it in its next cell. Helpers like this are how it grasps without a VLA policy.

def lower_guarded(arm_name, target_tcp_z):
    initial=robo.state().arm(arm_name)
    target=np.array(initial.eef_pos)
    final_eef_z=target_tcp_z-0.12*initial.approach[2]
    while target[2]>final_eef_z+0.001 and not robo.done:
        before=robo.state().arm(arm_name)
        target[2]=max(final_eef_z,target[2]-0.008)
        result=robo.move_to(arm_name,target,substeps=8)
        after=robo.state().arm(arm_name)
        print('descent',result.reached,tuple(round(v,4) for v in after.tcp_pos))
        if result.terminated or not result.planned or not result.reached or after.eef_pos[2]>=before.eef_pos[2]-0.001 or np.linalg.norm(np.array(after.eef_pos[:2])-target[:2])>0.01:
            print('stop',result)
            break
lower_guarded('right',0.800)
robo.show('right_wrist')
robo.show('head')

Cell 7 · RoboTwin 2.0, place object stand, seed 100002

It also chooses when to look: 84% of cells ask for a camera image, and 20% look only if the task is not yet done.

When things go wrong

Code helps most after a setback. PyRUA-Lean alone solved 96 task instances; tool calling alone solved 36. In most of the instances only code solved, both agents met the same setback, most often a VLA grasp that missed or an object that fell. Tool calling then asked the policy again, or stopped; the code agent closed the gripper itself, scripted the recovery from the same primitives, and checked each step. Two of them:

The dropped moka pot

LIBERO-PRO · “turn on the stove and put the moka pot on it”

Both agents turned the stove on and dropped the pot. Tool calling asked the VLA policy to grasp the fallen pot four times and ended the episode. The code agent grasped it itself, fitted a plane to the points of its lid to find how it leans, turned it upright, and set its base, not the gripper, over the burner.

Tool calling (RPent)gives up at LLM call 32
PyRUA-Leansolved at LLM call 33

Simulator recordings at 10× speed.

Camera views of both agents after the LLM calls named Excerpts of the code agent's cells

Top: the recorded camera views after the LLM calls named. Bottom: the code agent's cells, excerpts exactly as the model wrote them; “...” marks lines left out. The plane fit in call 23 raised an error (NumPy's linear algebra refuses half-precision arrays); call 24 cast the points and ran it again.

The cabinet door, pulled along its arc

RoboCasa365 atomic · “Open the cabinet door.”

Both agents located the door's hinge. Tool calling pulled the handle a few centimeters per LLM call and was still pulling when its 40 calls ran out. The code agent took the hinge as the centre of the circle through three positions of the handle, two of which it copied from earlier outputs, and pulled along that arc in one cell, stopping if a move fell short or the grip slipped. The door opened, and a VLA call in call 34 completed the task.

Tool calling (RPent)still pulling when its 40 LLM calls ran out
PyRUA-Leansolved at LLM call 34

Simulator recordings at 15× speed.

Camera views of both agents Excerpts of the code agent's cells

Top: the recorded camera views; bottom: the code agent's cells, excerpts exactly as the model wrote them.

Results

More tasks solved. Across 700 task instances, PyRUA-Lean raises the success rate from 63.1% to 71.7%, an improvement of 8.6 percentage points, or about 14% relative. It gains on every benchmark: 11.0 points on LIBERO-PRO, 9.2 on RoboTwin 2.0, 7.8 on RoboCasa365 atomic tasks and 5.0 on composite tasks.

Fewer calls and tokens for the same tasks. On the instances both agents solved, PyRUA-Lean reduces mean LLM calls from 17.0 to 8.7 and input tokens from 788k to 276k, reductions of 49% and 65%. At list prices, an episode costs $0.74 instead of $1.63; the saving in dollars is smaller than in tokens, in part because caching discounts the history that tool calling sends again and again.

The same success for a fraction of the tokens. Cut every recorded episode at a token budget and count what is solved by then: on LIBERO-PRO, PyRUA-Lean reaches tool calling's final success rate with 564k tokens per episode, tool calling with 3.18M.

Success rate against a token budget per episode on four benchmark groups

Success rate if every episode were stopped once its token consumption reached the budget on the horizontal axis. Dashed lines mark the budget at which each agent first reaches tool calling’s final success rate, with its value on the axis; the large number is their ratio. RC365: RoboCasa365.

Token-savings decomposition into fewer LLM calls and fewer tokens per call

Mostly from calling the model less often. The token ratio splits into the ratio of LLM calls and the ratio of input tokens per call. On LIBERO-PRO, the 4.48× token ratio comprises 2.55× fewer calls and 1.76× fewer tokens per call. On RoboCasa365 atomic tasks, PyRUA-Lean's calls are larger on average, yet fewer calls still reduce the total.

Also without VLA policies or guides. Removing the VLA policies, the operating guides or both from both agents, PyRUA-Lean keeps the higher success rate and the lower token use. Without VLA policies it reaches 68.8% on RoboTwin 2.0 against 42.8% for tool calling, composing classical primitives into approach, gripper control and state checks.

How we compared

LIBERO-PRO

40 tasks × 5 seeds = 200 instances

Tabletop tasks from its four perturbed suites. Franka arm; VLA policy π0.5.

RoboTwin 2.0

50 tasks × 5 seeds = 250 instances

Dual-arm tasks with randomized scenes. VLA policy LingBot-VLA.

RoboCasa365

18 atomic + 32 composite tasks × 5 seeds = 250 instances

Kitchen tasks for a mobile manipulator. VLA policy RLDX-1.

Both agents use GPT-6 Astra at high reasoning effort through the Codex CLI, with the same robot stacks, primitive implementations and VLA policies, and SAM 3 for segmentation. Each task instance runs once with each agent, under a budget of 40 LLM calls per episode (100 on the composite tasks), and each benchmark's own success check decides the outcome. The tool-calling baseline is RPent, whose robot stacks PyRUA-Lean builds on.

Try it

PyRUA-Lean drives the robot stacks of RPent. Install it into the Python environment of an RPent robot stack; the README covers the three robot stacks and a full episode.

git clone https://github.com/DAGroup-PKU/PyRUA-Lean.git && cd PyRUA-Lean
pip install -e .          # into the Python environment of your RPent robot stack
pyrualean api             # the robot API the agent sees; no simulator needed

BibTeX

@misc{si2026fewer,
  title={Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14\% Higher Success Rate but 65\% Fewer Tokens},
  author={Ruiyang Si and Jianxin Bi and Shunyu Yang and Rui Ni and Wenbo Huang and Qiang Wang and Shulong Jiang and Duomin Wang and Xiuyu Li and Haiwen Feng and Zhen Dong and Daquan Zhou},
  year={2026},
  eprint={2610.01939},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.01939}
}