Independent Ph.D. dissertation research, MIT Department of Urban Studies and Planning
Advisor: Andres Sevtsuk · Expected completion: June 2027
Seeing HAI is the title and organizing framework of my dissertation. It is distinct from Sidewalk Ballet, a City Form Lab collaboration and public installation that I lead at MIT. The two projects intersect, but they have different scopes, research ownership, and outputs.
Research question
Pedestrian counts describe how many people move through a street, but they do not show whether people stop, talk, rest, vend, play, or spend time together. Two sidewalks can carry similar foot traffic while supporting very different forms of public life. My dissertation asks how visual AI can extend close observation of these activities across cities without allowing a model’s available labels to define what matters.
The research combines ideas from urban studies with computer vision, vision-language models, and geospatial analysis. Its aim is to produce measures that remain interpretable as they move from individual street images to sidewalk segments and city-scale comparisons.
Data and framework
I assembled 5.7 million timestamped street-level panoramas covering almost 500 U.S. cities. The processing pipeline reconstructs sidewalk-facing views and codes each visible person across ten observable dimensions, including social grouping, posture, mobility state, activity, and spatial context.
The dissertation develops this work in three connected parts:
- Relational detection. Detect social groups as relationships among people rather than as a conventional object category.
- Measurement of sidewalk life. Extend the annotation framework from social grouping to a wider range of visible activities and interactions.
- Urban analysis and validation. Aggregate observations to sidewalk segments, compare them with pedestrian flow and street characteristics, and examine when the resulting measures are credible across places and data sources.
The Great Streets visualization
The Great Streets is a working interface for moving between citywide patterns and the street-level observations behind them. The prototype connects segment-level indicators in a 3D map with source street-view images and person-level activity and interaction labels. This makes the measurement pipeline inspectable: a user can select a street segment, compare observations across imagery sources and times, and trace an aggregate value back to the visible evidence.
Current research outputs
MINGLE
MINGLE is the first method developed within the dissertation. It combines person detection, depth-aware vision-language reasoning, and group aggregation to identify socially interacting groups in street imagery. The project uses 79,265 human judgments for model training and evaluation and releases annotations and metadata for 100,000 urban street images.
The paper was published in the AAAI 2026 AI for Social Impact Track:
Liu, L., Kudaeva, A., Cipriano, M., Al Ghannam, F., Tan, F., de Melo, G., & Sevtsuk, A. (2026). “MINGLE: VLMs for Semantically Complex Region Detection.” Proceedings of the AAAI Conference on Artificial Intelligence, 40(45). Paper
Ongoing dissertation work
The remaining work develops Human Activity and Interaction indicators at the sidewalk-segment level and evaluates their relationships with accessibility, pedestrian flow, street design, and neighborhood context. Because this part of the dissertation is ongoing, the website reports it as a research plan rather than as a completed empirical finding.