Research
MoveOD: Synthesizing Origin-Destination Commute Distribution from U.S. Census Data
Overview Research area: Intelligent transportation systems / cyber-physical systems data generation; specifically origin–destination (OD) travel demand synthesis from public census and geospatial data

- arXiv
- 2510.18858
- Published
- 2025-10-21
- Authors
- Rishav Sen, Jose Paolo Talusan, Abhishek Dubey, Ayan Mukhopadhyay, Samitha Samaranayake, Aron Laszka
AI summary
Overview
Research area: Intelligent transportation systems / cyber-physical systems data generation; specifically origin–destination (OD) travel demand synthesis from public census and geospatial data.
Technical level: Intermediate. The paper combines a Bayesian decomposition of a joint distribution with a constrained integer linear program, graph-based routing on road networks, and benchmark vehicle-routing experiments, but each component is explained at a level accessible to readers with a basic background in probability, optimization, or transportation data.
One-sentence scope: MoveOD is an open-source, automated pipeline that fuses five public data sources to produce building-level, minute-level commuter origin–destination flows for any U.S. county, demonstrated on Hamilton and Davidson counties in Tennessee.
What This Paper Is About
High-resolution origin–destination commute data is needed for traffic modeling, signal timing, congestion pricing, and vehicle routing, but outside a handful of data-rich cities such data is rarely available. The paper's goal is to synthesize a complete joint distribution over origin building, destination building, departure time, and arrival time by combining publicly available census marginals (LODES flows, ACS departure-time and travel-time tables) with building footprints and road networks, while enforcing that the synthetic trips reproduce the observed marginals.
Key Contributions
-
A modular synthesis pipeline (MoveOD) that generates OD commute datasets for any U.S. county and, in principle, for any region with comparable census datasets. The stated core contribution is the synthesis approach itself rather than a static dataset.
-
Integration of five open data sources into one OD dataset: ACS departure-time and travel-time distributions, LODES residence-to-workplace flows, county geometries (TIGER/Line block groups), road network information from OpenStreetMap, and building footprints from OSM and Microsoft Building Footprints (MSBF), plus an optional INRIX road-speed feed.
-
A constrained sampling plus integer-programming calibration method that reconciles the dataset with ACS and LODES through three anchors: matching commuter totals per origin zone, aligning workplace destinations with employment distributions, and calibrating travel durations to ACS-reported commute times.
-
An end-to-end automated system and demo: a browser dashboard requiring only a county and a year, demonstrated on Hamilton and Davidson counties in Tennessee and connected to a digital-twin simulator with a benchmark suite of classical and learning-based vehicle-routing algorithms.
Main Findings
-
Departure times reproduce the census exactly: Figure 1 overlays the ACS departure-time histogram (B08302) with the synthetic departures and the two curves coincide, confirming block-level departure times are reproduced exactly.
-
Calibration improves travel-time fidelity but not perfectly: Compared with the initial commuter assignment, the calibrated assignment aligns more closely with the ACS B08303 distribution in the left tail (up to 15 minutes) and the right tail (20 to 90 minutes) while preserving the mode and overall shape. However, the calibrated dataset slightly underestimates the proportion of typical travel times (15–19 minutes) compared to ACS, so discrepancies remain around the mean of the distribution.
-
Best hyperparameters: Sweeping α and β over [0,1] in increments of 0.1, α = 1 and β = 1 provide the smallest combined gap between the initial OD dataset slacks and the ACS travel-time distribution slacks.
-
Runtime scales linearly in trips: Overall runtime is O(M · log Y + Z) for M trips, Y network nodes, and Z road edges. MoveOD runs in 22 minutes for ~336K commuters (Davidson County, TN) and 14 minutes for ~150K commuters (Hamilton County, TN) on a 32-core, 5.2 GHz, 32 GB RAM Unix machine.
-
The generated instances are tractable and solvable: On a benchmark of n = 100 commuter trips sampled from Hamilton County early-morning departures (12:00 am to 4:59 am), with 30 vehicles of 5-passenger capacity and a 5-minute CPU cap, LKH-3 achieved the lowest vehicle miles traveled (2,255.16), followed by Clarke–Wright (2,622.84), Large Neighborhood Search (2,960.06), Genetic Algorithm (3,032.40), Simulated Annealing (3,239.05), Insertion Heuristic and Ant Colony Optimization (both 3,281.13), POMO (3,296.79), and the Attention Model (3,302.41). Passenger miles traveled was 1,455.74 for every algorithm, and coverage was 100 for all.
-
Consistency across algorithm families: Heuristics such as LKH-3 and Clarke–Wright produced near-optimal routes and learning-based methods generalized reasonably well to the held-out test region, supporting the practical usability of the synthetic demand. (Some entries in the reported Vehicle Utilization column are not cleanly rendered in the available text, so those values are not restated here.)
-
Tunable survey-bias correction: The pipeline exposes a mean road-speed shift factor ψ = τ̄_init / τ̄_ACS, justified by the observation that self-reported travel times usually overestimate actual travel times due to rounding errors and inclusion of parking or waiting times.
Methodology in Plain English
The authors target the joint probability P(O, D, S, E) — origin census unit, destination census unit, departure time, and arrival time — and factor it by the chain rule into P(E | S, D, O) · P(S | O) · P(D | O) · P(O). They assume departure-time blocks depend only on the origin.
Each factor comes from a different public source: P(O) from the set of census units in a county; P(D | O) from LODES residential-to-workplace flow counts; P(S | O) from ACS Table B08302 departure-time blocks; and P(E | S, D, O) from combining road-network travel times with ACS Table B08303 travel-time bins.
The generation proceeds in stages. First, buildings are selected in a tiered way: within each census unit, tagged residential OSM buildings are preferred for origins and tagged commercial OSM buildings for destinations; if none exist, Microsoft Building Footprint buildings are used; if neither exists, the census-unit centroid is used as a fallback. Second, each of the n_(o,d) commute trips is assigned an origin and destination building uniformly within the relevant census units. Third, a departure minute is sampled uniformly inside the chosen ACS departure-time block, which preserves the origin and departure-time marginals by construction. Fourth, a route is computed between the two buildings using graph distance and a time-dependent shortest-path travel time, giving an implied speed and an arrival minute.
Because raw ACS travel times are self-reported and biased high, the authors apply a multiplicative road-speed shift ψ to road-segment speeds so the mean synthetic travel time matches the ACS mean.
Finally, calibration is posed as an integer linear program per origin census unit. Decision variables are the calibrated number of commuters a_(o,d,s) for each destination and departure block, plus non-negative slacks: η for mismatches against the ACS travel-time bins and ζ for deviations from the initial counts. The objective minimizes a weighted sum of both sets of slacks. Three constraints enforce that total commuters per origin match N_o, that summed counts per destination equal N_o · p_(o,d), and that summed counts per departure block equal N_o · p_(o,s); a fourth constraint matches travel-time marginals up to slack; a fifth keeps calibrated counts close to the initial dataset. The calibrated counts are then mapped back to buildings by resampling and the arrival time is set using the mean-speed-shifted travel time.
For the routing benchmark, the authors draw n = 100 trips without replacement, treat each commuter as a pickup–delivery pair with unit demand, use the centroid of all census units as the depot, treat edge cost as great-circle distance between buildings (or driving distance/time if a road network is supplied), and give the pickup node the commuter's departure time. Training and test sets are split geographically: a contiguous 20% of census units is randomly designated as the test region, 100 trips are sampled from those origins, and all other census units form the training region. Every algorithm receives the same building coordinates, cost matrix, capacity of 5 passengers, and fleet of 30 identical vehicles.
Why This Matters
Impact on research. Most transportation cyber-physical systems research relies on datasets such as NYC taxi traces that do not reflect the target community, which limits the generalizability of models and control policies. MoveOD provides a community-representative demand layer that lets downstream work validate pipelines against local conditions instead of geographically mismatched proxies, and it does so for any U.S. county rather than a few metros.
Real-world applications:
- Predicting county-level traffic flows from where and when people actually commute.
- Public-transit design, including sizing fleets and scheduling services from finer-grained OD tables.
- On-demand and multimodal service planning, as demonstrated in the digital-twin dashboard with on-demand service zones and depots.
- Efficient road-network design and congestion pricing.
- Calibrating discrete-choice models and improving multi-modal assignment accuracy.
Industry relevance. The paper notes that over 3,000 smaller transit agencies across the United States lack integrated data pipelines. An automated tool that requires only a county and a year lowers the barrier for those agencies and for analysts who currently depend on proprietary surveys or traffic datasets such as INRIX that cannot be decomposed into individual OD trips. The digital twin can ingest GTFS feeds and on-demand configurations, letting operators switch between fixed-line, on-demand, and multimodal systems and compare performance.
Future Directions
-
Expand beyond weekday commuting. MoveOD currently generates OD data only for weekday commute trips; the authors plan to add weekend and non-commute trips.
-
Incorporate environmental and temporal variation. Accounting for weather and seasonal effects is listed as future work.
-
Ingest real-time traffic speed data. The framework already supports an optional INRIX feed and hybrid OSM/INRIX speeds; extending to real-time feeds is an explicit next step.
-
Generalize to other data sources and regions. The synthesis approach is designed to ingest any macro-level movement dataset and enhance spatial resolution to building level, and can be adapted to other countries with comparable census datasets — including integration with crowd-sourced GPS trajectories or emerging IoT mobility streams.
Target Audience
Transportation researchers and practitioners who need granular OD demand but lack proprietary survey data; CPS and reinforcement-learning researchers who need community-representative, time-resolved demand for closed-loop simulation; transit agency analysts and urban planners evaluating fixed-line, on-demand, or multimodal service designs; and graduate students or engineers interested in data fusion, constrained sampling, and integer-programming-based calibration applied to urban mobility.
Authors’ abstract
High-resolution origin-destination (OD) tables are essential for a wide spectrum of transportation applications, from modeling traffic and signal timing optimization to congestion pricing and vehicle routing. However, outside a handful of data rich cities, such data is rarely available. We introduce MOVEOD, an open-source pipeline that synthesizes public data into commuter OD flows with fine-grained spatial and temporal departure times for any county in the United States. MOVEOD combines five open data sources: American Community Survey (ACS) departure time and travel time distributions, Longitudinal Employer-Household Dynamics (LODES) residence-to-workplace flows, county geometries, road network information from OpenStreetMap (OSM), and building footprints from OSM and Microsoft, into a single OD dataset. We use a constrained sampling and integer-programming method to reconcile the OD dataset with data from ACS and LODES. Our approach involves: (1) matching commuter totals per origin zone, (2) aligning workplace destinations with employment distributions, and (3) calibrating travel durations to ACS-reported commute times. This ensures the OD data accurately reflects commuting patterns. We demonstrate the framework on Hamilton County, Tennessee, where we generate roughly 150,000 synthetic trips in minutes, which we feed into a benchmark suite of classical and learning-based vehicle-routing algorithms. The MOVEOD pipeline is an end-to-end automated system, enabling users to easily apply it across the United States by giving only a county and a year; and it can be adapted to other countries with comparable census datasets. The source code and a lightweight browser interface are publicly available.