Design judgment + human-centered AI

Making AI tools that sharpen design judgment.

I build and study interfaces for an age of generative abundance: tools that help people explore alternatives, notice consequential differences, explain tradeoffs, and build taste over time.

Cognitive Science Ph.D. student at UC San Diego Advised by Prof. Steven P. Dow in ProtoLab.

Thesis thread Scaffolding for Taste
Methods Build, test, iterate

Thesis thread

Scaffolding taste in an age of generative abundance.

Scaffolding for Taste asks a simple question: when AI can make ten plausible options in a minute, how do people learn to notice the difference that matters? I study interfaces that support reflective practice: making a move, seeing what changed, comparing alternatives, and building better reasons for choosing one direction over another.

Design principle

Design help should not just make more artifacts. It should make the next move easier to see, and the next judgment easier to explain.

Don Norman's Design Lab talk helped sharpen this framing: the next generation of designers needs tools that cultivate taste, not just tools that generate options.

I am testing this frame across creative tools, learning environments, mixed reality, everyday robotics, and community or civic design settings: places where tools can help people notice viable options, weigh tradeoffs, and choose the next move.

Why now

Why this work matters now

AI makes it easy to produce variants; my question is how people learn to choose well.

01

More output does not mean more taste.

When generation is cheap, the scarce work is structuring the design space: what dimensions matter, what to vary, what to preserve, and how to justify a direction.

02

Everyone is closer to the prototype.

Designers, researchers, PMs, and engineers all shape early artifacts now. Interfaces need to support shared reasoning before decisions harden.

03

Speed still needs reflection.

Fast cycles are useful only if teams keep room for critique, incubation, and explaining why a direction is better.

Research focus

Design, Evaluate, Situate.

I keep testing the same loop: how to make alternatives comparable, let evidence sharpen the next question, and fit assistance to the situation where judgment happens.

Showing the Design motion sketch. Vary the alternatives, compare them, then refine.

Design lens

Human-AI for design thinking

I study interfaces that support design thinking beyond artifact generation: dimensional scaffolds that make the design space visible, variation mechanisms that expose alternatives, structured comparison that helps people reason across options, and staged fidelity that lets ideas mature from rough intent to defensible direction.

See DesignWeaver
Moves
  1. 01

    Externalize useful dimensions so people can see what matters.

  2. 02

    Use variation and structured comparison to reason across alternatives.

  3. 03

    Move from coarse ideas to committed artifact choices through staged fidelity.

Exploring

Adjacent contexts for the same loop

These threads are places I test the same concern with judgment, agency, and evidence in context.

Situate Everyday robotics Integrating intelligence into robots so they can better assist people in everyday living and working contexts. Design Learner-centered AI Scalable support that keeps agency and the learning process visible. Evaluate Post-deployment iteration Using traces from real use to improve benefit and catch harm. Situate Community-centered tools Coordination and sensemaking for groups with shared ownership.

Motion note Inspired by motion craft from Stripe's design team and Katie Dill at Stripe Sessions. Interaction rules grounded by Jeffrey Heer and collaborators on animated transitions and CMU's Data Interaction Group .

Live field The visible Dot Orbit layer uses Paper Shaders by Paper; the restraint also follows the founder's design walkthrough.

Selected publications

Papers worth opening first.

  1. Short form Workshop First author 0 cites
    Sirui Tao, William P. McCarthy, and Steven P. Dow
    In Herding CATs: Making Sense of Creative Activity Traces (CHI 2026 Workshop), Apr 2026
    Workshop Position Paper
    When this paper helps. This paper helps when you can see what someone did in an AI tool but still cannot tell what that action meant to them.

    This paper helps when you can see what someone did in an AI tool but still cannot tell what that action meant to them.

    Look for
    • You are designing lightweight feedback at the moment a creative workflow becomes ambiguous.
    • You need a unit of analysis that joins a short trace, interface state, and a user's optional explanation.
    It adds

    A proposal for trace-guided micro-episodes and a utility-for-rationale pattern that turns useful recovery controls into chances to learn why something happened.

    Its limit

    It is a position paper and research agenda, not an empirical validation of the interventions or their downstream effects.

    I checked this against the paper and its official record on 2026-07-13.

    Read the full evidence and citation context
  2. CVPR
    hotspot.png
    Full paper Coauthor 18 cites
    Zimo Wang, Cheng Wang, Taiki Yoshino, Sirui Tao, Ziyang Fu, and Tzu-Mao Li
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun 2025
    Highlight
    When this paper helps. This paper helps when an implicit-surface argument needs more than the usual eikonal condition—especially around convergence, stability, topology, or surface area.

    This paper helps when an implicit-surface argument needs more than the usual eikonal condition—especially around convergence, stability, topology, or surface area.

    Look for
    • You are comparing objectives for reconstructing signed distance functions from unoriented points.
    • You need theoretical and experimental evidence about a screened-Poisson alternative.
    It adds

    A heat-loss objective with convergence and stability analysis, plus 2D and 3D reconstruction evidence and a natural surface-area penalty.

    Its limit

    The reported gains belong to the evaluated reconstruction settings; sparse boundaries and parameter choices can still produce failure modes.

    I checked this against the paper and its official record on 2026-07-13.

    Read the full evidence and citation context
  3. CHI
    designweaver.png
    Full paper First author 41 cites
    Sirui Tao, Ivan Liang, Cindy Peng, Zhiqing Wang, Srishti Palani, and Steven P. Dow
    In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Apr 2025
    When this paper helps. This paper helps when a blank prompt box is asking novices to supply design vocabulary they have not had a chance to learn.

    This paper helps when a blank prompt box is asking novices to supply design vocabulary they have not had a chance to learn.

    Look for
    • You are designing dimensional scaffolds for text-to-image exploration.
    • You want evidence about prompt language, visual diversity, novelty, and rising expectations.
    It adds

    A palette that externalizes product-design dimensions, grounded in expert practice and evaluated with 52 novice designers.

    Its limit

    The study used one chair-design task with novices, and richer prompts did not reliably improve satisfaction or expectation alignment.

    I checked this against the paper and its official record on 2026-07-13.

    Read the full evidence and citation context
  4. NeurIPS
    physion.gif
    Full paper Coauthor 187 cites
    Daniel Bear, Elias Wang, Damian Mrowca, Felix Binder, Hsiao-Yu Tung, Pramod RT, Cameron Holdaway, Sirui Tao, Kevin Smith, Fan-Yun Sun, Fei-Fei Li, Nancy Kanwisher, Josh Tenenbaum, Dan Yamins, and Judith Fan
    In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Dec 2021
    When this paper helps. This paper helps when you need a reproducible benchmark for asking whether a model predicts physical outcomes in human-like ways.

    This paper helps when you need a reproducible benchmark for asking whether a model predicts physical outcomes in human-like ways.

    Look for
    • You are comparing object-centric and non-object-centric visual models with human judgments.
    • You need a common contact-prediction task across several simulated physical scenarios.
    It adds

    Eight scenario families, matched human and model evaluation, and released data and code for systematic comparison.

    Its limit

    Synthetic scenes and binary contact prediction do not establish broad real-world physical understanding.

    I checked this against the paper and its official record on 2026-07-13.

    Read the full evidence and citation context
Open full publication list

Updates

Recent signals.

Full timeline

HotSpot got selected as a CVPR 25 Highlight!

Research opportunities

For students who like messy questions.

I like working with curious, motivated, and kind undergraduate and master's students, especially people who have a question they cannot stop poking at.

You do not need to show up as a polished researcher. It helps if you like reading carefully, making small prototypes, testing claims, looking honestly at evidence, and writing clearly about what changed.

The best fit is someone with real stake in a domain or problem, plus enough patience to turn that interest into a concrete study.

Reading closely Prototyping Evaluation Analysis Communication

Start here

  1. Read the research-start post.
  2. Email s1tao@ucsd.edu with subject "UCSD Research Interest".
  3. Include a 1-page CV or resume, an optional portfolio link, and a 3-5 sentence note on the questions or domains you care about.

Connect

Find me around the web.

Advisor Lineage & Past Collaborations

Ph.D. Steven P. Dow HCI
Master's Steven P. Dow HCI Tzu-Mao Li Graphics
Undergrad Judith E. Fan Cognition & Intuitive Physics