A live web app that turns the camera, mic, and touchscreen you already own into short tests of reaction time, hand coordination, and fine motor control, then tracks how you change over time.
Baseline started with one question for the quarter: what happens when I use AI to turn everyday devices into tools that test, teach, and track human physical and cognitive function? My answer was a browser platform with no special equipment. You hold your hands up to your laptop camera to test coordination, use your fingertip to test fine motor precision, and make common hand signs to test reaction time. Every result gets saved, so the point isn't a single score. It's watching your own numbers move over weeks.
I'm a designer, not an engineer, and I built the whole thing myself, frontend, backend, and deployment, with AI writing most of the code. That made my real job setting AI up to work inside my decisions, and catching the moments when its easy answer was the wrong one. (It had a lot of easy answers!)
Before I let AI write anything, I set the project up so it would work inside my decisions instead of making its own. My first session had no AI in it at all. I built the folder structure and file setup by hand so I'd always know where everything lived and what each file was for.
I wrote a scope markdown file and a PRD covering what Baseline measures, who it's for, and that it's not diagnostic. AI worked from those every session, and any suggestion that drifted got checked against them.
GAME_STANDARD.md spelled out the structure every test had to follow. When AI's output drifted, I fed the rules back instead of patching things by hand. After losing a session of work once, I committed to git frequently. I was not letting that happen again!
Every time I pushed back on or reshaped an AI suggestion, I logged what it suggested, what I did instead, and why. I also had it explain code before I kept any of it.
That log became my Records of Resistance, 54 entries across 20 sessions. I kept receipts! Here are a few examples:
The rest are in the full Records of Resistance.
I started by looking at existing assessment tools, including ACE-X, the NIH Toolbox, and José Azuara's work, to understand what these tests usually measure and who they're built for. I kept my research notes raw, typos and half-formed questions included, because cleaning them up would have erased the uncertainty that was driving the project.
That research turned into a position statement I wrote before any code existed. Four lines, and every decision after this point got checked against them:
That last line was the one I treated as non-negotiable. A tool that reads your body through a webcam could easily slide into making health claims it has no business making, and it was on me to hold that line when feature ideas started pushing past it.
Early on, the AI recommended I focus on one polished game. I decided to be more ambitious instead, even on a short timeline, and built a multi-test platform, choosing to cut other scope to make room for it. That one decision created everything after it: the homepage hub, a shared test structure, a backend, and a history page.
Then I spent the rest of the project protecting that scope by cutting things:
The design work that mattered most was about measurement. The first version of a test usually produced a clean number that quietly measured the wrong thing, and for an assessment tool that's worse than a bug, because nothing looks broken. Per-hand scoring was the first of these calls. The others:
My first few tests were each built differently, and I spent a lot of time manually correcting the AI to make them feel consistent. GAME_STANDARD fixed that for users as much as for the AI. Every test now follows the same flow:
After that, new tests came out coherent almost immediately. When the AI's first version of Rhythm Sync skipped the practice round, I fed it the standard and it fell back in line. Users learn the flow once and it works the same on every test. Muscle memory for the win!
The visual side got the same treatment. I rejected the emoji icons the AI first gave the test cards, because they made the platform feel like a game, and research I'd read suggested heavily gamified health tools can feel too fun to trust. I replaced them with a dark blue, white, and light blue system that's professional without being clinical.
I put the platform in front of two outside testers, ages 21 and 58. Both validated the core idea and both asked the same thing: they wanted to see their results as a graph. That pushed tracking over time from a feature to the center of the product. I added Chart.js graphs with a toggle between total score and hand symmetry. Testing also surfaced a bug where red dots silently couldn't be collected, which turned out to be a JavaScript quirk treating an ID of 0 as empty. (Zero is a number, JavaScript!)
On the backend, I used Supabase for accounts and saved results, and deployed the whole platform to Vercel myself.
Taking the time to understand the backend first (Record #29) paid off at the end. A security review from my professor flagged three issues, including one where my data looked protected but wasn't: the access rules were comparing mismatched data types, so they never actually ran. Very confident lock, no door. Because I knew where the data lived and how the rules worked, I found it quickly. The fixes:
I also chose not to add password complexity rules yet. For a v1 with no sensitive medical data, it was friction without much benefit, and I'd rather make that call on purpose than by default.
The biggest thing I learned about the product took the whole project to understand:
What matters isn't a single score. What matters is seeing how your abilities change over time. That's the meaningful data. Not the number, but the arc.
From my artist statement
The biggest thing I learned about working with AI is that it's a tool I can steer. Its suggestions often sounded reasonable, and slowing down to ask whether they actually served the user is how I learned which ones were worth keeping. Going forward, I understand how to use AI as a tool in my UX workflow without outsourcing my decision-making to it. I know how to push back on suggestions, how to ask clarifying questions about the thinking behind a proposal, and how to stay grounded in user needs when AI offers something that feels easy but doesn't actually serve the design.
If I started over, I'd write GAME_STANDARD before building the first test instead of after the fourth. Setting the framework first feels slower, but it compounds into speed later. I'd also test with more people. Two testers were enough to validate the concept, but not enough to say the tests are reliable.
Every session, decision, and override is documented in the full process blog, along with the rest of the process book.