Semantically Adaptive Control Barrier Functions with Vision-Language Models for Real-Time Robot Navigation

A vision-language model reads the robot's camera view and adjusts how cautious its safety controller is, without blocking the control loop.

Simulation and physical Clearpath Jackal results for AlphaAdj.

Overview

Abstract

Robots operating in dynamic, unstructured environments must maintain safety without becoming unnecessarily conservative. While control barrier functions (CBFs) provide a principled mechanism for collision avoidance, their behavior is typically determined from geometric state alone and therefore cannot distinguish between semantically different situations with similar physical configurations. We present AlphaAdj, a vision-to-control navigation framework that uses egocentric RGB observations and a vision-language model (VLM) to adapt CBF conservativeness according to visual semantic context. The VLM produces a scalar contextual caution signal that dynamically modulates the CBF parameter α, while a geometry- and speed-dependent cap bounds admissible aggressiveness. To support real-time operation despite VLM latency, inference runs asynchronously with the control loop, and the semantic signal is extracted directly from token-level output probabilities rather than full autoregressive text generation. We evaluate AlphaAdj across matched-geometry, dynamic-human, and cluttered navigation scenarios, as well as on a physical Clearpath Jackal. When obstacle geometry is held fixed, AlphaAdj allocates 0.405 m more clearance to a human than to a benign object, while geometry-only and non-visual adaptive CBF baselines produce no semantic differentiation. This additional behavioral reserve improves robustness to unexpected human motion, including cluttered conditions in which all evaluated geometric baselines collide while AlphaAdj remains collision-free. The system maintains a 0.0167 s control period with 0.158 s mean VLM round-trip latency, demonstrating that semantic visual reasoning can complement geometric safety control without entering the critical control path.

Why AlphaAdj

Geometry sets the limit. Vision decides how much of it to use.

1

More caution for people

AlphaAdj gives a person more clearance than a physically identical object in the same spot.

2

Caution carries through the maneuver

If a person starts moving after the robot has already assessed the scene, that caution carries into the avoidance instead of resetting.

3

Inference doesn't block control

The VLM runs in the background while the 60 Hz control loop keeps executing on the most recent valid estimate.

Method

From camera frame to control parameter

RGB observation → caution score → α → CBF-QP
AlphaAdj system diagram. Every 5 control steps, an asynchronous branch captures an RGB frame, queries the VLM, and passes the result through a persistence gate to produce the applied caution score. Every control step, a geometric branch observes robot state, computes the geometry and speed cap, and fuses it with the caution score through a stale gate to set alpha for the CBF-QP. A timeline below shows five overlapping VLM calls against a continuously ticking 60 Hz control loop.

Two loops run at different rates. A slow, asynchronous branch turns a camera frame into a caution score every 5 control steps. A fast branch runs every step, computing the geometry and speed cap and fusing it with the latest caution score into α for the CBF-QP.

Persistence gate

The caution score stays tied to whichever obstacle is currently most relevant to the CBF, even if it leaves the camera view during an avoidance maneuver.

Fusion and stale gate

A geometry- and speed-based cap bounds α at every step. A late or missing VLM result falls back to that cap, so it can never make the robot less safe than geometry alone.

Results

Results

+0.405 m
more clearance for a human than a physically identical object
+1.4%
path overhead in the matched benign-object case
Collision‑free
Across the tested dynamic-human trials
60 Hz
control-loop rate, independent of the 0.158 s mean VLM round trip
Matched-geometry test

Same geometry, different clearance

Robot start, goal, and obstacle position stay identical. Only the visible identity of the target changes.

Matched-geometry comparison. Geometry-only, OA-CBF, and PACBF produce identical trajectories regardless of target identity. AlphaAdj gives the human more clearance than the object, at a small path cost compared to a globally conservative fixed-alpha controller.
Experiment videos

Recorded runs

Four simulation runs and three physical Clearpath Jackal runs, shown at real relative timing.

Simulation

IsaacSim, Carter differential-drive robot

Only the target's identity changes between runs. The person gets more clearance than the box.

A person crosses the robot's path at different speeds. Across part of that range the geometry-only controller collides and AlphaAdj does not.

The person starts stationary, then moves without warning once the robot is already navigating. The robot avoids contact.

Same unexpected motion, now with added clutter. This is the only method here that avoids contact.

Physical Jackal

Clearpath Jackal, onboard camera

Onboard the Jackal, the caution score changes as the camera sees different objects and people.

The Jackal gives a wider berth to a person than to a chair in the same spot.

The Jackal avoids a person crossing its path.

Supplementary material

Prompt formulation

A single prompt produces the caution score. Several phrasings were tested before choosing this one.

Final prompt

“A mobile robot shares this space with whatever is visible ahead. How strongly should the robot yield and avoid imposing on, disturbing or crowding what it sees? 0 = no need to yield, 1 = yield strongly. Return strict JSON only: {"score": <one decimal between 0.00 and 1.00>}.”

AlphaAdj reads the score from the model's next-token probabilities instead of waiting for generated text. One forward pass gives a probability over the digits 0 to 9, and the score is their weighted average. This keeps the signal continuous and avoids the delay of full text generation.

Contextual caution scores across six scenes: cardboard box 0.357, pallet 0.388, cart 0.470, a single human 0.513, two humans 0.567, a group spread across the robot's path 0.607.

The score is graded, not a binary human/object switch. It responds to broader scene context, including group size and layout.

Prompt strategyIntended behaviorObserved effectSelection
General yield framing Ask how strongly the robot should yield to whatever is visible, without naming any object class. Broad, graded response across object types and crowd sizes, beyond a simple person or object switch. Selected
Consequence framing Ask about the consequence of a close interaction with the scene. Reduced overall variation. Compressed the difference between people and ordinary objects. Reduced generality
Sensitivity framing Ask about how sensitive or vulnerable the scene is. Stronger response to some non-human clutter than to people present. Reduced person sensitivity
Person-explicit framing Ask directly about proximity to people. The sharpest person-versus-object separation of any formulation tested. Excluded, too narrow for a general-purpose signal

The selected framing stays general. It responds to movable equipment and groups of people, not only a fixed “human” category, and stays graded rather than acting as a hard switch.

Supplementary material

Score-to-alpha mapping and tuning

The caution score sets α through a normalized sigmoid. Its shape and center point were chosen by comparing several alternatives.

The final k8_c45 sigmoid mapping. Left, the normalized caution score as a function of raw caution score, with a steep transition centered at 0.45. Right, the resulting alpha for near and far obstacles, bounded below by alpha_min = 0.10.

The final mapping, k8_c45: a sigmoid with slope k = 8 centered at c = 0.45, normalized to [0, 1] and used to set α between the geometry cap and a floor of 0.10.

MappingBehaviorEvaluation takeawayStatus
Shallow sigmoid (k=2) Gradual semantic response, centered at the midpoint of the score range. Smooth, but slower to react to a clearly cautious scene. Evaluated
Steep sigmoid, center 0.50 (k8_c50) Sharp response centered at 0.50. Strong differentiation across most tested conditions. Evaluated
Steep sigmoid, center 0.45 (k8_c45) Sharp response, center shifted slightly earlier. Earlier caution response. Selected final balance across the evaluated conditions. Selected

A single, globally conservative fixed α can reach a similar clearance around a person, but only by applying that same caution to every object it encounters, at roughly 3.5× the path cost of AlphaAdj's selective reserve in the benign case (see the matched-geometry figure above).