TEMPEST, a 2.8MB AI model, hits 91.71% accuracy on a 45 driver dataset and only loses 4.3 points as the pool grows. The privacy question is what the paper does not ask.
A 2.8MB research model called TEMPEST, described in an arXiv preprint, turns 60 seconds of driving into a stable identifier. The system maps a minute of accelerator, brake, steering, and motion signals to a 96-number summary, then matches that summary against an enrolled driver list. It trains in 50 epochs, and adding a new driver does not require retraining. The system updates its enrollment table in place.
On a 45-driver dataset, the model reaches 91.71% Rank-1 accuracy, meaning the correct driver is the system's top guess more than nine times out of ten. That beats the best classical baseline in the comparison by 17.9 percentage points and the strongest triplet-loss system by 58.4 points, according to the abstract.
The more interesting number is what happens as the driver pool grows. Existing systems in this family degrade quickly. Supervised triplet-loss pipelines drop 22 percentage points when the enrollment set expands from 10 to 45 drivers; unsupervised variants drop 32.5 points. TEMPEST loses 4.3 points across the same expansion. On the public KIA Soul dataset, the same model posts a 7.3-point gain over the best classical approach within a single recording session and a 14.3-point gain across separate sessions. The cross-session test, where a driver's behavior is matched days later under different conditions, is the harder one.
The mechanism is the part that generalizes. Older driver-ID pipelines leaned on triplet loss, a training setup that pulls similar drivers close and pushes different drivers apart, but only one trio at a time. Triplet loss overfits to the patterns of any one recording session and fragments as the pool grows. TEMPEST swaps that for an additive angular margin loss borrowed from face recognition, called ArcFace, which forces every driver into a distinct slice of a normalized angular space. Each 60-second window becomes a 96-dimensional vector, a short list of 96 numbers that summarizes who was behind the wheel, and the geometry of that space is what keeps the model from collapsing as new drivers are added.
The HTML version of the paper confirms that the embedding is small enough to be checked at the edge, on a car computer, without round-tripping to a cloud service. That size is the structural change. A 2.8 MB model ships in a fleet telematics unit, an insurance dongle, or a rental-car gateway. It also runs invisibly.
The paper is silent on the consumer and regulatory side. There is no discussion of consent, no mention of who controls the enrollment database, and no analysis of what a fleet operator, insurer, or law enforcement request would do to the model. The KIA Soul dataset used for cross-session testing is a public corpus collected by a Korean security research group; the driver list is a closed, consented set. The deployment case is not.
A passive identifier built from 60 seconds of driving expands agency if consent and transparency are designed in. A driver who enrolls themselves, who can see their own record, and who can delete it has more control over a shared vehicle than a key fob alone provides. The same model erodes agency if it is bolted on by an insurer or a fleet operator who never asks. The technical answer to "is this accurate?" is now well-established. The design answer to "is this fair?" is not.
The result is a preprint, not a peer-reviewed paper. Author names, affiliations, and submission date are not stated in the hydrated abstract, and the comparison baselines are described only by category rather than by name. Replication and benchmark design will matter; the 45-driver test is a controlled set, not a population. The structural claim, that angular-margin training turns a temporal embedding into a graceful-scaling identifier, is the part most likely to outlast the specific result. It also travels: any short behavioral signal with consistent structure, from typing cadence to gait, fits the same pipeline.
The model is open to read at arXiv:2610.06855. What gets built on top of it is the open question.