Research / GPL-TR-2026-01
The Shape of Identity
Modeling behavioral continuity through time, context, and drift.
How much of behavioral identity survives time, context, and practice?
- Author
- GrayPass
- Version
- GPL-TR-2026-01
- Date
- 7 May 2026
Finding
Behavior changes across time and context.
Behavioral identity is a moving distribution, not a fingerprint.
People type and move in ways that are measurably their own, but the pattern is not a fingerprint. It practices, tires, and drifts. On two familiar benchmarks, the same detector on the same people reports nearly twice the error once the evaluation is forced to respect time. The constructive conclusion is a change of object: identity as a per-person distribution with trait, context, and drift components, and a seven-item reporting standard for anyone making longitudinal behavioral-identity claims.
Visual analysis
01 / Model
Identity is a moving distribution.
Person-specific structure can persist while context and collection time move the observations around it.
FIG. 1Trait, context, and drift
Conceptual model · no data plotted
Trait
People occupy different distributions.
Context
A changed setting moves observations together.
Drift
One person moves across collection time.
02 / Trait structure
Traits and states coexist.
Hold times retain moderate person-specific structure across eight sessions. Inter-key latencies stay below the same threshold.
FIG. 2Intraclass correlation across eight sessions
CMU keystroke benchmark
Hold times
11 featuresCMU keystroke benchmarkKeydown to keydown
10 featuresCMU keystroke benchmarkKeyup to keydown
10 featuresCMU keystroke benchmarkRange across the 11 hold-time features
Inter-key latency features below 0.500
03 / Practice
The population does not sit still.
The same people type faster across the study, shifting the distribution before any detector changes.
FIG. 3Practice moves the population
CMU keystroke benchmark
Median duration in session 8 relative to session 1
Subjects faster in session 8 than at enrollment
04 / Protocol and drift
The time split changes the answer.
With the detector and data held fixed, respecting collection order nearly doubles error. The farther verification moves from enrollment, the higher the measured error.
FIG. 4One detector, three time splits
CMU benchmark · scaled Manhattan detector
Session-disjoint
Enrollment uses session 1 and verification uses sessions 2 to 8.
95% CI 15.2% to 22.0%
50 enrollment repetitionsCMU keystroke benchmarkProtocol penalty from same-session to session-disjoint
FIG. 5Error rises with distance from enrollment
CMU benchmark · session-1 enrollment
- Session-disjoint EER
- Endpoint 95% interval
A same-session result measures separation at one moment. Longitudinal claims need time-disjoint evaluation and a drift curve.
Method
Read the abstract and method summary
Behavioral signals from ordinary computer use, such as typing rhythm and pointer movement, carry measurable person-specific structure, and systems that treat this structure as an identity signal often report strong laboratory accuracy. Reported accuracy, however, depends on an easily overlooked choice: how the evaluation splits data across time. We reanalyze two canonical benchmarks, the CMU keystroke dynamics benchmark (51 subjects, 8 sessions) and the Balabit mouse dynamics challenge (10 users), holding the detector fixed and varying only the evaluation protocol. Hold-time features show intraclass correlations of 0.55 to 0.77 across eight sessions, while all inter-key latencies fall below 0.5; median typing duration shortens by 26% over the eight sessions, and 96% of subjects type faster in the final session than in the first. With a scaled Manhattan detector and 50 template repetitions, mean equal error rate moves from 10.0% (same-session) to 14.3% (session-mixed) to 18.6% (session-disjoint) on identical data, and error grows with session distance from enrollment, from 14.6% to 21.7%. A small mouse-dynamics cohort replicates the pattern of strong trait structure alongside protocol-sensitive verification error. We propose a distributional decomposition of behavioral identity into trait, context, and drift components, and a reporting standard for longitudinal behavioral-identity claims.
- A quantified evaluation-protocol penalty on canonical benchmarks, with subject-bootstrap confidence intervals throughout.
- A trait-versus-drift decomposition for each feature family: hold times are trait-like with moderate drift; latencies are weakly personal with large practice-driven drift.
- A cross-modality replication on the Balabit mouse cohort, reported with the wide intervals a ten-user population requires.
- A distributional framing of behavioral identity separating trait, context, and drift, and what identity evidence should mean for continuous verification.
- A compact seven-item reporting standard for longitudinal behavioral-identity claims.
Evidence and sources
CMU keystroke dynamics benchmark
51 subjects typing the same strong password 400 times over 8 sessions; 20,400 repetitions; 31 timing features.
Trait structure, practice nonstationarity, and the three-protocol verification experiment.
Balabit Mouse Dynamics Challenge
10 users, 65 remote-desktop sessions, 2,253,816 pointer events, roughly 177 hours of wall-clock time.
Small cross-modality replication, stated as direction, not magnitude.
Limitations
- The CMU benchmark is fixed-text password entry by a university population on laboratory apparatus; absolute error levels and the practice curve are specific to that setting.
- The benchmark provides session order and a one-day minimum separation, not calendar dates, so no claim is made in units of elapsed time.
- The Balabit cohort has ten users, uncontrolled tasks, and unverified session order; its protocol contrast is indicative only.
- Detectors are deliberately simple and untuned; stronger detectors would lower absolute error but cannot remove the underlying distributional movement.
- Practice, fatigue, motivation, and time of day are confounded in session order; no causal claim is made among them.
Not claimed
- No claim that any behavioral signal uniquely identifies a person. Intraclass correlations well below 1 say the opposite.
- No performance of any product or deployed system. Every EER characterizes a benchmark under a stated protocol with a deliberately simple detector.
- No universal accuracy figure for keystroke or mouse biometrics, and the report should not be cited as one.
Reproducibility
Everything is deterministic under seed 42. A single script regenerates every number from the source data into a machine-readable results file, and every figure from the same file.
Citation
GrayPass (2026). The Shape of Identity: Modeling Behavioral Continuity through Time, Context, and Drift. GPL-TR-2026-01. https://www.graypass.org/research/the-shape-of-identity
Continue the work
Read the complete paper, then follow the next question.
The PDF contains the complete method, secondary results, limitations, references, and publication record.