Skip to content
AI.info

Research

Context and Symmetry in Auditing: A Case Study of Skeleton Inference in Motion Capture

Overview Research area: AI auditing, AI safety and ethics, with a case study drawn from motion capture (mocap) and anthropometric measurement. Technical level: Intermediate. The core argument is conce

arXiv
2608.10194
Published
2026-08-10
Authors
Emma Harvey, Emanuel Moss, Hauke Sandhaus, Abigail Z. Jacobs, Mona Sloane

AI summary

Overview

Research area: AI auditing, AI safety and ethics, with a case study drawn from motion capture (mocap) and anthropometric measurement.

Technical level: Intermediate. The core argument is conceptual and qualitative, but the case study assumes familiarity with motion capture pipelines, body segment parameters, and concurrent validity testing.

One-sentence scope: The paper introduces "contextual audits" as a method for auditing AI measurement systems within the practices that produce their outputs, and argues that the science and technology studies concept of "symmetry" can enable audits when ground truth is unknown, unknowable, or contested.

What This Paper Is About

AI audits typically compare a system's actual outputs against nominal expected outputs, but this requires deciding what inputs are relevant and what "correct" behavior looks like. For AI systems that measure human bodies, no agreed-upon ground truth often exists — even anatomical measurements such as forearm length can be defined in multiple, incompatible ways. The paper proposes contextual auditing, a method that grounds the audit in real measurement practices and treats competing measurement modalities symmetrically rather than crowning one as ground truth, illustrated with a case study of skeleton inference in motion capture.

Key Contributions

  1. Contextual audits as a method. The paper introduces the contextual audit, defined as an audit that makes claims about the accuracy of system outputs within the context of the practices that produce them. It shows how this framing enables auditors to interrogate what serves as ground truth and to identify the consequences of using any particular ground truth to evaluate a system.

  2. Symmetry as a tool for audits without ground truth. The authors adapt the STS principle of symmetry — treating competing theories of truth with the same analytical approach, as if either might be true or false — to AI auditing, so that auditors can treat each measurement modality as provisionally true and examine the resulting tensions and assumptions.

  3. A demonstration through case study. The paper applies contextual auditing to skeleton inference in motion capture, focusing on the context of mocap for creative design, and produces a completed "Contextual Auditing Matrix" comparing a marker-based mocap system against tape measure-based anthropometry.

  4. A method adaptation and agenda. The authors adapt the socio-technical matrix from Sloane et al. (2022), extended by Rhea et al. (2022b) and Sloane et al. (2025), to measurement systems, and propose future directions for auditing motion capture systems.

Main Findings

  • Contextual audits require qualitative research first. The authors argue that standard audits, which use pre-existing or synthetic data, fail to capture real-world use. A contextual audit should instead include field observation or ethnography of users, informational interviews, and a review of research, data, design decisions, and system documentation. They point to prior work (Pruss 2023 on a recidivism prediction tool; Cheng et al. 2022 on a child welfare assessment tool) as evidence that audits divorced from actual use can be meaningless or can overstate bias.

  • Practitioners treat validity as an article of faith. In fieldwork with mocap practitioners at NYU, the researchers observed that practitioners attended to camera calibration, marker placement, removal of reflective objects from the stage, and awareness of areas of the stage with greater and lesser accuracy. Yet practitioners judged measurements by whether they "look[ed] right" and recalibrated when measurements "g[o]t wonky" (elongated necks, twisted limbs, foreshortened arms). In the absence of egregious distortions, they trusted the calibrated system rather than verifying measurements against an alternate apparatus.

  • Ground truth for body segment parameters is contested. The paper argues there is no single agreed-upon system of ground truthing for BSPs. For example, shoulder distance could be measured as collarbone length or as the straight-line distance between shoulder blades; it could be based on bones alone (not directly observable inside the body) or include soft tissue. The authors state that universal ground truth values for BSPs are unknown and potentially unknowable, and that early ground truth data on BSPs used in motion capture relied on dismembering dead bodies or implanting pins into the bones of living bodies.

  • Anthropometry encodes its own assumptions. The paper notes that anthropometric units of measurement were originally socially constructed (e.g., based on easy-to-measure lengths such as armspan) and later standardized, and that anthropometric landmarks were originally chosen to study variation in physical anthropology and later to aid biometric identification in criminal justice and for eugenic purposes (Maguire 2009). The matrix lists the assumption that people have standard body landmarks and that landmarks are equally identifiable regardless of body shape or size, with the anticipated harm that individuals far from the anthropometric average may have eugenicist harms reified or reinforced.

  • Mocap systems encode assumptions about whose bodies count. The paper states that modern mocap systems and their inferences are largely built on data from a small number of thin male cadavers (Harvey et al. 2024), and that prior work (Durkin and Dowling 2003) found mocap inferences do not generalize equally well to individuals whose bodies do not match those in the data. The authors' matrix lists the assumption that people have standard body landmarks corresponding to those of primarily male research subjects and a standard amount of soft tissue between markers and the body surface (Keller et al. 2023), with the expected harm of less faithful representation for people who are not male or not thin.

  • Harms can travel from entertainment to higher-stakes settings. Although mocap for creative design is relatively low-stakes, the authors argue it still carries representational harms if it misrepresents or erases a body type in entertainment, and that materials, competences, and meanings developed for entertainment are often adapted for higher-stakes settings. They cite the Microsoft Kinect, originally marketed as a video game controller, becoming the basis for home health and rehabilitation mocap applications due to its low price and small size.

  • Two hypotheses were formulated, but the empirical results are not reported in the provided content. The authors state they expected to see less fidelity of representation for people who are not male or not thin, and less fidelity of representation for people who are not close to the average used to construct anthropometric measurement procedures and norms. The paper content supplied here ends mid-sentence in the audit protocol description, before any quantitative results, so no accuracy, stability, or comparison figures are available.

  • No numerical audit results are reported in the provided content. Benchmark accuracies, effect sizes, and statistical comparisons across the facets of body size, sex, and time are not present in the available text.

Methodology in Plain English

The authors combine conceptual argument with a hands-on case study.

Conceptual work. They first review how AI auditing is normally done and identify its two hard questions: which inputs to use, and what a system's outputs should ideally look like. They then import two ideas from sociology and STS. From social practice theory, they take the framing that practices are made of materials, competences, and meanings that shift over time. From STS, they take symmetry, the principle of analyzing competing claims as if each could be true or false rather than assuming one is the standard. Applied to AI auditing, symmetry means treating each measurement modality as provisionally true and examining where the modalities disagree.

The matrix. They adapt an existing socio-technical matrix into a Contextual Auditing Matrix that auditors fill in before running any quantitative tests, by answering five questions: (1) what objects are captured and inferred; (2) how those objects are made legible to the system; (3) how ground truth is and historically has been established; (4) what assumptions those "ground truths" encode; and (5) how those assumptions may produce harms in the relevant context of use. They complete the matrix for two side-by-side modalities: marker-based motion capture and tape measure-based anthropometry.

The fieldwork. They worked with mocap practitioners at NYU whose work covers augmented, virtual, and extended reality and external clients including companies, museums, and artists. The site uses a 40'x35' mocap stage surrounded by an OptiTrack Prime 13 system with 24 cameras. The authors spent a day in training — calibrating cameras, identifying body landmarks and affixing markers, capturing motion, and troubleshooting — then conducted two non-consecutive weeks of field experiments in June and July 2023, recruiting 24 participants. They also conducted informal interviews with practitioners and took extensive field notes. The study was approved by the NYU Institutional Review Board (IRB-FY2023-7677).

The audit protocol. The protocol had two phases. In Phase 1, body measurements were taken by pairs of researchers using a soft tape measure and, for weight, a digital scale, following National Health and Nutrition Examination Survey protocols. Measurements recorded, in order, were: standing height, wingspan, sitting height, upper leg length, biacromial breadth, upper arm length, upper arm circumference, abdominal circumference, hip circumference, thigh circumference, head circumference, and weight. Most measurements were read aloud by a primary measurer and confirmed and recorded by an assistant. In Phase 2, measurements were recorded with the mocap system. Participants wore a form-fitting spandex mocap suit with a Velcro exterior. Daily room calibration used the OptiTrack large space calibration wand moved in a figure-8 pattern until at least 10,000 samples were registered per camera, with the team proceeding only when the system reported an "Exceptional" rating (approximate mean error ≤ 0.6 mm). The ground plane was defined with an L-shaped calibration square aligned with participants' walking direction. The team attached 41 retroreflective markers following the standard OptiTrack "Baseline 41" body model, with periodic quality control checks by a facility mocap practitioner. Participants assumed a static T-pose for subject calibration, with neither manual skeleton alignment nor dynamic calibration performed. Participants then followed a scripted protocol of everyday movements — T-pose, walking, sitting, rotating, squatting, arm swings, arm circles, jumping, and kicking — before the provided text cuts off.

The analytical stance. Rather than using the tape measure as ground truth, the authors compare how the two modalities' measurements vary with respect to one another, without characterizing differences as systematic bias from a ground truth.

Why This Matters

The paper reframes a fundamental problem in AI auditing: many systems measure things for which no verifiable ground truth exists, and pretending otherwise can hide harms. By making assumptions explicit before quantitative testing, the approach is intended to prevent auditors from silently importing the assumptions of whichever measurement modality they happen to treat as correct.

Real-world applications:

  • Workplace safety monitoring, where mocap-derived measurements of workers' bodies inform health and safety decisions, and where error rates across body types could translate into unequal protection.
  • Gait recognition and biometric surveillance, which relies on body segment parameters and risks differential performance across the bodies it is applied to.
  • Physical therapy and rehabilitation, including low-cost consumer mocap hardware repurposed from gaming, as in the Microsoft Kinect case.
  • Entertainment and creative design, where mocap is used to record and distort bodies, and where misrepresentation or erasure of body types constitutes a representational harm even at low physical stakes.

Industry relevance. The paper speaks directly to mocap software and hardware vendors (the case study examines OptiTrack and its Motive software and body models), to studios and facilities that produce mocap content, and to organizations building or procuring AI systems that measure people. It suggests that documentation and historical design decisions embedded in commercial body models are legitimate and necessary objects of audit, and that auditors should follow local site practices rather than academic best practices if they want ecologically valid results.

Future Directions

  • Auditing higher-stakes mocap contexts. The authors propose that the areas their case study flags should be scrutinized in future audits of workplace safety monitoring, biometric surveillance, and medical diagnosis, given that mocap materials and practices migrate from entertainment into these settings.

  • Completing and reporting the empirical comparison. The case study protocol describes testing the stability of body segment parameters across the facets of body size, sex, and time, but the provided content stops before any results. Reporting how the mocap and tape measure modalities vary relative to one another across these facets is the obvious next step the paper sets up.

  • Testing symmetry as an auditing practice. The paper positions symmetry as a practical tool for pointing to where assumptions should be investigated further, but leaves open how auditors should act on the tensions it surfaces — whether to justify one modality over another, modify a measurement practice, or supplement the audit with additional techniques.

  • Extending contextual auditing beyond motion capture. The authors present the Contextual Auditing Matrix as adapted for measurement systems generally, raising the question of how well it transfers to other domains where ground truth is contested, such as creditworthiness or toxicity benchmarking, which the paper cites as analogous cases.

Target Audience

AI auditors and audit researchers; AI ethics and AI safety scholars; science and technology studies researchers interested in applied measurement; motion capture practitioners, studios, and tool vendors; and product, policy, and compliance teams responsible for systems that measure human bodies. The paper is most useful to readers who want a methodological framework for auditing systems where no definitive ground truth exists, rather than readers looking for quantitative performance benchmarks.

Authors’ abstract

Humans are increasingly expected to interact with AI systems that observe and make inferences about them - but do these systems actually work? A standard approach to answering this question is AI auditing. Conducting an AI audit requires identifying how a system behaves (i.e., determining what types of inputs to audit it with and then observing and documenting actual system behavior) and contrasting that with how a system should behave (i.e., determining what the nominal outputs of a system should look like). We argue that this is best done through a contextual audit, which we introduce as a method for auditing measurements within the context of the practices that produce them. We show how contextual auditing enables the interrogation of assumptions implicit in the audit process and allows auditors to be explicit about what serves as ground truth, which we define as verifiable measurements about the real world against which systems are evaluated. We outline how the concept of symmetry from science and technology studies can enable audits when ground truth is unknown, unknowable, or contested. Finally, to demonstrate how contextual and symmetric audits can be conducted in practice, we present a case study of skeleton inference in motion capture and propose areas suggested by our case study as particularly fruitful for future audits of motion capture systems.

Read the original paper