←
AI 201 Capstone, solo and shipped

Baseline

A live web app that turns the camera, mic, and touchscreen you already own into short tests of reaction time, hand coordination, and fine motor control, then tracks how you change over time.

Role Designer and builder (solo)
Built with Claude, Google AI Studio, MediaPipe, Supabase, Vercel
Timeline 10 weeks, 20 sessions
54Records of Resistance, times I pushed back on AI and logged why
3Sensor-based tests live on the platform
51Values patched after a security review
20Sessions documented, from first prototype to deploy
OVERVIEW

Baseline started with one question for the quarter: what happens when I use AI to turn everyday devices into tools that test, teach, and track human physical and cognitive function? My answer was a browser platform with no special equipment. You hold your hands up to your laptop camera to test coordination, use your fingertip to test fine motor precision, and make common hand signs to test reaction time. Every result gets saved, so the point isn't a single score. It's watching your own numbers move over weeks.

I'm a designer, not an engineer, and I built the whole thing myself, frontend, backend, and deployment, with AI writing most of the code. That made my real job setting AI up to work inside my decisions, and catching the moments when its easy answer was the wrong one. (It had a lot of easy answers!)

HOW I WORKED WITH AI

Set the rules first, then let AI build

Before I let AI write anything, I set the project up so it would work inside my decisions instead of making its own. My first session had no AI in it at all. I built the folder structure and file setup by hand so I'd always know where everything lived and what each file was for.

A scope file it read first

I wrote a scope markdown file and a PRD covering what Baseline measures, who it's for, and that it's not diagnostic. AI worked from those every session, and any suggestion that drifted got checked against them.

Written rules to follow

GAME_STANDARD.md spelled out the structure every test had to follow. When AI's output drifted, I fed the rules back instead of patching things by hand. After losing a session of work once, I committed to git frequently. I was not letting that happen again!

Resisting on purpose

Every time I pushed back on or reshaped an AI suggestion, I logged what it suggested, what I did instead, and why. I also had it explain code before I kept any of it.

That log became my Records of Resistance, 54 entries across 20 sessions. I kept receipts! Here are a few examples:

#2

A platform, not one game

AI suggested
Focus on one strong, polished prototype.
I chose
A multi-test platform with a shared test structure, a backend, and a history page.
Why
The whole point was tracking how different abilities change over time. One game couldn't do that, no matter how polished it was.
#19

Per-hand scoring instead of a total

AI suggested
One total score for the Bilateral Coordination Test: count the dots, regardless of which hand grabbed them.
I chose
Each hand scored separately, with symmetry as the headline number: the weaker hand's score divided by the stronger hand's. Left 24, right 23 comes out to 96%.
Why
With a total, someone using only their right hand scores the same as someone using both hands equally. For a bilateral test, that hides the only thing worth knowing.
#29

Understanding before building

AI suggested
A menu of backend options (Firebase, Supabase, Node/Express) to pick from.
I chose
Not to pick yet. I asked what a backend actually is, learned how the frontend, API layer, and database connect, and then chose Supabase.
Why
I didn't want to build on something I couldn't explain. That understanding is exactly what let me find a hidden security bug at the end of the project.
#54

Writing my own reflection

AI suggested
A polished draft of my artist statement.
I chose
To throw it out entirely and write my own, built from three questions I answered myself.
Why
The draft was well written, but it was Claude's voice, not mine. A reflection in someone else's words isn't a reflection.

The rest are in the full Records of Resistance.

01 / RESEARCH

Deciding what Baseline would never be

I started by looking at existing assessment tools, including ACE-X, the NIH Toolbox, and José Azuara's work, to understand what these tests usually measure and who they're built for. I kept my research notes raw, typos and half-formed questions included, because cleaning them up would have erased the uncertainty that was driving the project.

That research turned into a position statement I wrote before any code existed. Four lines, and every decision after this point got checked against them:

That last line was the one I treated as non-negotiable. A tool that reads your body through a webcam could easily slide into making health claims it has no business making, and it was on me to hold that line when feature ideas started pushing past it.

02 / SCOPE

Bigger in one direction, smaller in every other

Early on, the AI recommended I focus on one polished game. I decided to be more ambitious instead, even on a short timeline, and built a multi-test platform, choosing to cut other scope to make room for it. That one decision created everything after it: the homepage hub, a shared test structure, a backend, and a history page.

Then I spent the rest of the project protecting that scope by cutting things:

03 / MEASUREMENT DESIGN

Making sure each test measured what it claimed to

The design work that mattered most was about measurement. The first version of a test usually produced a clean number that quietly measured the wrong thing, and for an assessment tool that's worse than a bug, because nothing looks broken. Per-hand scoring was the first of these calls. The others:

04 / STANDARDS

One flow, learned once

My first few tests were each built differently, and I spent a lot of time manually correcting the AI to make them feel consistent. GAME_STANDARD fixed that for users as much as for the AI. Every test now follows the same flow:

After that, new tests came out coherent almost immediately. When the AI's first version of Rhythm Sync skipped the practice round, I fed it the standard and it fell back in line. Users learn the flow once and it works the same on every test. Muscle memory for the win!

The visual side got the same treatment. I rejected the emoji icons the AI first gave the test cards, because they made the platform feel like a game, and research I'd read suggested heavily gamified health tools can feel too fun to trust. I replaced them with a dark blue, white, and light blue system that's professional without being clinical.

05 / TESTING & SHIPPING

Real people, a real backend, a real security review

I put the platform in front of two outside testers, ages 21 and 58. Both validated the core idea and both asked the same thing: they wanted to see their results as a graph. That pushed tracking over time from a feature to the center of the product. I added Chart.js graphs with a toggle between total score and hand symmetry. Testing also surfaced a bug where red dots silently couldn't be collected, which turned out to be a JavaScript quirk treating an ID of 0 as empty. (Zero is a number, JavaScript!)

On the backend, I used Supabase for accounts and saved results, and deployed the whole platform to Vercel myself.

Taking the time to understand the backend first (Record #29) paid off at the end. A security review from my professor flagged three issues, including one where my data looked protected but wasn't: the access rules were comparing mismatched data types, so they never actually ran. Very confident lock, no door. Because I knew where the data lived and how the rules worked, I found it quickly. The fixes:

I also chose not to add password complexity rules yet. For a v1 with no sensitive medical data, it was friction without much benefit, and I'd rather make that call on purpose than by default.

06 / REFLECTION

What I'm taking with me

The biggest thing I learned about the product took the whole project to understand:

What matters isn't a single score. What matters is seeing how your abilities change over time. That's the meaningful data. Not the number, but the arc.

From my artist statement

The biggest thing I learned about working with AI is that it's a tool I can steer. Its suggestions often sounded reasonable, and slowing down to ask whether they actually served the user is how I learned which ones were worth keeping. Going forward, I understand how to use AI as a tool in my UX workflow without outsourcing my decision-making to it. I know how to push back on suggestions, how to ask clarifying questions about the thinking behind a proposal, and how to stay grounded in user needs when AI offers something that feels easy but doesn't actually serve the design.

If I started over, I'd write GAME_STANDARD before building the first test instead of after the fourth. Setting the framework first feels slower, but it compounds into speed later. I'd also test with more people. Two testers were enough to validate the concept, but not enough to say the tests are reliable.

Every session, decision, and override is documented in the full process blog, along with the rest of the process book.

← Back to portfolio Next project →