PHYSICAL AI · WORLD MODELS · SAFE AI

Learning from experience. Acting in the physical world.

Seeking PhD opportunities in Physical AI, world models, and safe embodied robotics. Get in touch · View CV

I am Bowen Jing (荆博闻), Tech Lead at Tuojing Intelligence and a master’s graduate of the University of Manchester. My research connects Physical AI, World Models, and Safe AI in embodied robotics: understanding physical interactions, anticipating action consequences, and making safety part of successful task execution.

I approach these questions through the design of data distributions. Within the observed distribution, policies should preserve and express meaningful behavioral preferences rather than average them away. Beyond its coverage, I pursue targeted generation of rare, safety-critical scenarios for learning and evaluation. The question is which experiences a robot needs—and how to construct and evaluate them.

My methods include vision-language-action models (VLA) and world-action models (WAM) for connecting perception, prediction, and control, alongside diffusion and flow matching for generative modeling and controlled sampling. These methods support learning from behavioral diversity and constructing experience beyond routine data collection.

Reproducing a trajectory is only one part of reproducing an interaction. My Real2Sim / Sim2Real agenda connects environment properties with robot dynamics, considering approach velocity, hand preshaping, contact timing and sequence, and the resulting forces, deformation, and slip. Vision and touch ground this work in physical feedback.

Long term, I want robots to pursue goals, learn from accumulated experience, and adapt through interaction, enabling safe participation in everyday human life.

Selected Research

Three connected questions: which behavioral differences should a policy preserve, what makes an interaction acceptable, and how can we test situations that recorded experience rarely covers?

StyleDrive research overviewAAAI 2026 · Oral

01 / Preserve meaningful behavioral differences

StyleDrive

When several actions are valid, whose preference should a policy follow?

Second author

Driving demonstrations do not prescribe a single response: different drivers can make different choices in the same situation. Treating these differences as noise risks learning a default behavior that obscures individual preferences. StyleDrive makes preference an explicit part of the learning problem.

Approach. Style-aware annotation, preference-conditioned policies, and the SM-PDMS metric bring behavioral diversity into both learning and evaluation.

Significance. This extends the question from whether a policy can drive competently to whether it can do so in a way that reflects a specified preference.

A dataset and benchmark for personalized end-to-end driving, with an explicit measure of driving-style alignment.

Project Page / Code • arXiv

Paper title & authors

StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
Ruiyang Hao, Bowen Jing, Haibao Yu, Zaiqing Nie

SoftVTBench research overviewECCV 2026 Workshop · Oral

02 / Define physically acceptable success

SoftVTBench

Is a task successful if the object is damaged along the way?

First author

For deformable objects, reaching the goal does not establish that the interaction was acceptable. A robot may finish the task while excessively deforming the object. SoftVTBench treats the physical consequences of contact as part of the problem definition, making this distinction measurable.

Approach. Synchronized visual, tactile, and action data are paired with deformation-aware tasks and FEM-based evaluation.

Significance. The benchmark makes physical constraints part of what a policy is evaluated against, and provides a basis for studying when touch helps distinguish task completion from acceptable interaction.

4,000 demonstrations · 40 tasks · 50+ deformable assets. FEM-based metrics evaluate deformation alongside task success.

Project Page • Tuojing Page • arXiv

Paper title & authors

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Bowen Jing*, Mingxin Wang*, Ruiyang Hao, Chenchen Ge, Hanwen Shen, Junjie He, Yang Cui, Yiming Hou, Weitao Zhou, Jiawei Wang, Minglei Li, Dandan Zhang, Ding Zhao, Houde Liu, Xiaofan Li, Si Liu, Ping Luo, Haibao Yu

CounterScene research overviewECCV 2026 Workshop · Oral

03 / Probe the limits of recorded experience

CounterScene

How can we expose failures that routine experience rarely reveals?

First author

A policy can perform well on recorded driving data while remaining vulnerable to rare interactions that those logs barely cover. CounterScene turns scenario generation into a targeted search for plausible conditions that expose these vulnerabilities.

Approach. Counterfactual causal reasoning and guidance during generation shape multi-agent behavior toward safety-critical interactions, which are evaluated in closed loop.

Significance. This gives the world model a role beyond reproducing typical behavior: it becomes a tool for actively examining the limits of a policy’s experience.

Closed-loop evaluation tests collision-inducing scenarios and transfer from Waymo Open Motion to nuPlan.

Project Page • arXiv

Paper title & authors

CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
Bowen Jing, Ruiyang Hao, Weitao Zhou, Haibao Yu

News

  • 2026.09:  🏆 Early versions of SoftVTBench and CounterScene were accepted as Orals at the Safe World Models for Trustworthy Embodied AI workshop, ECCV 2026.
  • 2026.08:  🪨 Our paper “KnockGS: Interaction-Grounded Calibration of Physical Gaussian Representations” is now on arXiv! Read here · Code
  • 2026.07:  🤖 Our paper “ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts” is now on arXiv! Read here
  • 2026.07:  🧈 Released SoftVTBench, a deformation-aware visuo-tactile dataset and benchmark for deformable-object manipulation. arXiv · Project Page
  • 2026.03:  🌍 Our paper “CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation” is now on arXiv! Read here
  • 2026.03:  🏗️ Our paper “ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction” is now on arXiv! Read here
  • 2025.11:  🏆 Our work StyleDrive has been accepted as an Oral presentation at AAAI 2026.

Additional Research

KnockGS: Interaction-Grounded Calibration of Physical Gaussian Representations
Chenchen Ge*, Hanwen Shen*, Bowen Jing, Jiyuan Cai, Xiaofeng Wang, Hongsen Lei, Weitao Zhou, Dandan Zhang, Haibao Yu

Calibrating elasticity and density from an object’s response to an applied force, then testing the estimated properties on held-out interactions.

Project Page • arXiv • Code

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li

World-action modeling for manipulation under visual distribution shifts. Reported zero-shot LIBERO-Plus improvement of 21.3 points over Fast-WAM, with real-world success under visual shift increasing from 25.8% to 61.5%.

arXiv

ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction
Haibao Yu, Kuntao Xiao, Jiahang Wang, Ruiyang Hao, Yuxin Huang, Guoran Hu, Haifang Qin, Bowen Jing, Yuntian Bo, Ping Luo

Feed-forward 4D Gaussian scene reconstruction for autonomous driving, combining static and dynamic representations for efficient novel-view synthesis.

Project Page • arXiv

Multi-modal Sensor Fusion for End-to-End Autonomous Driving
Bowen Jing — University of Manchester, MSc dissertation

My MSc dissertation investigated camera–LiDAR fusion with channel attention and GRU waypoint prediction, evaluated closed-loop in CARLA across urban scenes and weather conditions.

Dissertation

Experience

  • 2025.08 – Present · Tuojing Intelligence — Tech Lead, Real2Sim2Real and tactile simulation
  • 2025.02 – 2025.08 · Tsinghua University, AIR — Research Intern in Large-Scale Autonomous Driving Data Mining

Education

  • 2023.09 – 2024.09, MSc in Advanced Computer Science, University of Manchester
    • Focus: Deep Learning, Computer Vision, Reinforcement Learning, Robotics
    • Dissertation: Multi-modal Sensor Fusion for End-to-End Autonomous Driving
    • Graduated with Distinction (Top 10%)
  • 2020.09 – 2023.06, BSc in Computer Science, University of Manchester
    • Specialized in software development and machine learning foundations
    • Final Year Project: Spiking Neural Network