A multi institution digital twin of a quantum sensor shows that optimizing for one performance metric can quietly blow the other two, by up to two orders of magnitude.
Quantum sensors are sold on a single number: sensitivity, the smallest magnetic field the device can resolve. A new digital-twin study finds that optimizing for that number alone can wreck the other two metrics a sensor needs, by as much as two orders of magnitude.
The result comes from a multi-institution preprint that ran a full physics simulation of two sensor platforms: nitrogen-vacancy (NV) diamond ensembles and cesium optically pumped magnetometers (OPMs). At a fixed sensitivity target of 100 picotesla per square-root-hertz (pT/√Hz), the recovered-field bias on the digital twin ranged from 8 to 1,500 nanotesla (nT) depending on how the sensor handled temperature drift. The same setup that beat the noise floor on paper could, on the bench, misread a magnetic field by a factor of nearly 200.
Datasheets rarely quote the other two metrics, which is why so many quantum-sensor projects pass lab benchmarks and miss field targets. The new paper, published on arXiv as 2608.28519 and hydrated by Quantum Computing Report, argues that three metrics move together, and that optimizing one without modeling the others is how sensor builders quietly ship the wrong device.
The team, MITRE, Quantum Brilliance, NVIDIA, and SandboxAQ, built a single simulation pipeline that propagates a sensor's full quantum state through time, then attributes the error budget across the three metrics using a method borrowed from cooperative game theory. NVIDIA's cuQuantum library, specifically its cuDensityMat module, ran the math on graphics processing units (GPUs), which is what made the multi-parameter sweep tractable. Without the GPU acceleration, a single parameter combination could take hours; with it, the team could sweep the design space in minutes.
For NV diamond ensembles, the dominant limiter on sensitivity is dephasing, the loss of quantum coherence over time, which alone accounts for roughly 89 percent of the sensitivity budget. Systematic accuracy is capped by a temperature-dependent shift in the diamond's internal energy levels. Drift robustness, the metric engineers care about most for field-deployed sensors, is gated by leakage in the acousto-optic modulator that drives the optical readout. None of these terms is a single number; each couples to temperature, drive power, and operating frequency in different ways.
For cesium OPMs, the platform behind the clinical use case in the paper, the same digital twin found a separate set of attributions. The atomic spin-projection floor sits at about 1.1 femtotesla per square-root-hertz; the system-level floor rises to roughly 4.4 pT/√Hz after software common-mode rejection, a technique that subtracts noise shared across channels. The clinical target is 5 millimeters of dipole-localization error, the precision needed to localize the source of an abnormal heart rhythm inside a patient's chest.
The clinical target is the link between the simulation and SandboxAQ's magnetocardiography program and the running Mayo Clinic PRISM trial, an independent validation path first reported by FierceHealthcare. The digital twin does not yet model the patient, and the cesium OPM results in the paper are simulation outputs, not measured clinical data. The trial is where the trade-off curve has to hold up.
The most important caveat is the simplest one. arXiv:2608.28519 is a preprint, not peer-reviewed. The 8 to 1,500 nT bias range, the 89 percent dephasing share, and the 5-millimeter clinical target are all simulation results on a digital twin, useful for design, not yet a record of what a deployed sensor will do. The team's contribution is a way to find the real limiter before the hardware ships, not a claim that the limiters are now known.
Three things will tell if the trade-off curve holds: peer-review status of the preprint, independent replication of the curve on physical hardware, and whether the Mayo PRISM trial reports the sensitivity-and-bias split the digital twin predicts. The 100 pT/√Hz story only becomes a clinical story when both halves, simulation and patient, line up.