PHYSICAL AI · WORLD MODELS · SAFE AI
Learning from experience. Acting in the physical world.
Seeking PhD opportunities in Physical AI, world models, and safe embodied robotics. Get in touch · View CV
I am Bowen Jing (荆博闻), Tech Lead at Tuojing Intelligence and a master’s graduate of the University of Manchester. My research connects Physical AI, World Models, and Safe AI in embodied robotics: understanding physical interactions, anticipating action consequences, and making safety part of successful task execution.
I approach these questions through the design of data distributions. Within the observed distribution, policies should preserve and express meaningful behavioral preferences rather than average them away. Beyond its coverage, I pursue targeted generation of rare, safety-critical scenarios for learning and evaluation. The question is which experiences a robot needs—and how to construct and evaluate them.
My methods include vision-language-action models (VLA) and world-action models (WAM) for connecting perception, prediction, and control, alongside diffusion and flow matching for generative modeling and controlled sampling. These methods support learning from behavioral diversity and constructing experience beyond routine data collection.
Reproducing a trajectory is only one part of reproducing an interaction. My Real2Sim / Sim2Real agenda connects environment properties with robot dynamics, considering approach velocity, hand preshaping, contact timing and sequence, and the resulting forces, deformation, and slip. Vision and touch ground this work in physical feedback.
Long term, I want robots to pursue goals, learn from accumulated experience, and adapt through interaction, enabling safe participation in everyday human life.
Selected Research
Three connected questions: which behavioral differences should a policy preserve, what makes an interaction acceptable, and how can we test situations that recorded experience rarely covers?
01 / Preserve meaningful behavioral differences
StyleDrive
When several actions are valid, whose preference should a policy follow?
Second author
Driving demonstrations do not prescribe a single response: different drivers can make different choices in the same situation. Treating these differences as noise risks learning a default behavior that obscures individual preferences. StyleDrive makes preference an explicit part of the learning problem.
Approach. Style-aware annotation, preference-conditioned policies, and the SM-PDMS metric bring behavioral diversity into both learning and evaluation.
Significance. This extends the question from whether a policy can drive competently to whether it can do so in a way that reflects a specified preference.
A dataset and benchmark for personalized end-to-end driving, with an explicit measure of driving-style alignment.
Paper title & authors
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
Ruiyang Hao, Bowen Jing, Haibao Yu, Zaiqing Nie
02 / Define physically acceptable success
SoftVTBench
Is a task successful if the object is damaged along the way?
First author
For deformable objects, reaching the goal does not establish that the interaction was acceptable. A robot may finish the task while excessively deforming the object. SoftVTBench treats the physical consequences of contact as part of the problem definition, making this distinction measurable.
Approach. Synchronized visual, tactile, and action data are paired with deformation-aware tasks and FEM-based evaluation.
Significance. The benchmark makes physical constraints part of what a policy is evaluated against, and provides a basis for studying when touch helps distinguish task completion from acceptable interaction.
4,000 demonstrations · 40 tasks · 50+ deformable assets. FEM-based metrics evaluate deformation alongside task success.
Project Page • Tuojing Page • arXiv
Paper title & authors
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Bowen Jing*, Mingxin Wang*, Ruiyang Hao, Chenchen Ge, Hanwen Shen, Junjie He, Yang Cui, Yiming Hou, Weitao Zhou, Jiawei Wang, Minglei Li, Dandan Zhang, Ding Zhao, Houde Liu, Xiaofan Li, Si Liu, Ping Luo, Haibao Yu
03 / Probe the limits of recorded experience
CounterScene
How can we expose failures that routine experience rarely reveals?
First author
A policy can perform well on recorded driving data while remaining vulnerable to rare interactions that those logs barely cover. CounterScene turns scenario generation into a targeted search for plausible conditions that expose these vulnerabilities.
Approach. Counterfactual causal reasoning and guidance during generation shape multi-agent behavior toward safety-critical interactions, which are evaluated in closed loop.
Significance. This gives the world model a role beyond reproducing typical behavior: it becomes a tool for actively examining the limits of a policy’s experience.
Closed-loop evaluation tests collision-inducing scenarios and transfer from Waymo Open Motion to nuPlan.
Paper title & authors
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
Bowen Jing, Ruiyang Hao, Weitao Zhou, Haibao Yu
News
- 2026.09: 🏆 Early versions of SoftVTBench and CounterScene were accepted as Orals at the Safe World Models for Trustworthy Embodied AI workshop, ECCV 2026.
- 2026.08: 🪨 Our paper “KnockGS: Interaction-Grounded Calibration of Physical Gaussian Representations” is now on arXiv! Read here · Code
- 2026.07: 🤖 Our paper “ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts” is now on arXiv! Read here
- 2026.07: 🧈 Released SoftVTBench, a deformation-aware visuo-tactile dataset and benchmark for deformable-object manipulation. arXiv · Project Page
- 2026.03: 🌍 Our paper “CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation” is now on arXiv! Read here
- 2026.03: 🏗️ Our paper “ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction” is now on arXiv! Read here
- 2025.11: 🏆 Our work StyleDrive has been accepted as an Oral presentation at AAAI 2026.
Additional Research
KnockGS: Interaction-Grounded Calibration of Physical Gaussian Representations
Chenchen Ge*, Hanwen Shen*, Bowen Jing, Jiyuan Cai, Xiaofeng Wang, Hongsen Lei, Weitao Zhou, Dandan Zhang, Haibao Yu
Calibrating elasticity and density from an object’s response to an applied force, then testing the estimated properties on held-out interactions.
Project Page • arXiv • Code
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li
World-action modeling for manipulation under visual distribution shifts. Reported zero-shot LIBERO-Plus improvement of 21.3 points over Fast-WAM, with real-world success under visual shift increasing from 25.8% to 61.5%.
ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction
Haibao Yu, Kuntao Xiao, Jiahang Wang, Ruiyang Hao, Yuxin Huang, Guoran Hu, Haifang Qin, Bowen Jing, Yuntian Bo, Ping Luo
Feed-forward 4D Gaussian scene reconstruction for autonomous driving, combining static and dynamic representations for efficient novel-view synthesis.
Multi-modal Sensor Fusion for End-to-End Autonomous Driving
Bowen Jing — University of Manchester, MSc dissertation
My MSc dissertation investigated camera–LiDAR fusion with channel attention and GRU waypoint prediction, evaluated closed-loop in CARLA across urban scenes and weather conditions.
Experience
- 2025.08 – Present · Tuojing Intelligence — Tech Lead, Real2Sim2Real and tactile simulation
- 2025.02 – 2025.08 · Tsinghua University, AIR — Research Intern in Large-Scale Autonomous Driving Data Mining
Education
- 2023.09 – 2024.09, MSc in Advanced Computer Science, University of Manchester
- Focus: Deep Learning, Computer Vision, Reinforcement Learning, Robotics
- Dissertation: Multi-modal Sensor Fusion for End-to-End Autonomous Driving
- Graduated with Distinction (Top 10%)
- 2020.09 – 2023.06, BSc in Computer Science, University of Manchester
- Specialized in software development and machine learning foundations
- Final Year Project: Spiking Neural Network


