Every time I rejected an AI suggestion, refined a recommendation, or made a deliberate creative decision — I documented it. These records track my thinking across 6 milestones and 20 sessions, organized by project phase.
Sessions 2–3: Establishing quarter question clarity and project direction.
Rejected continuing VibeWord Animator as a prototype direction. The concept was technically interesting (Gemini-synthesized unique animations based on word vibes) and visually striking, but explored aesthetic AI generation rather than assessment or function measurement. This established the pattern: check alignment with quarter question *first*, before technical elegance.
Rejected the recommendation to focus on a single strong prototype and instead chose to build a multi-test platform. The quarter question asks what happens when AI turns everyday devices into *tools* (plural), and a single game wouldn't address tracking over time. This foundational decision shaped every subsequent choice: the need for a homepage hub, a common game standard, a backend for storing results across test types, and a history page showing multiple dimensions of performance.
Sessions 4–6: Understanding tools, developing research foundation, protecting core scope.
Flagged accepting AI's tool breakdown without developing my own understanding as a dependency risk. Rather than accepting the explanation, I committed to understanding *why* certain tools were chosen before making future technology decisions. The key distinction learned: `@mediapipe/hands` (raw 21-point coordinates for custom logic) versus `@mediapipe/tasks-vision` GestureRecognizer (named gesture labels). Understanding the difference became foundational to designing the multi-test platform.
Deliberately chose Replit as the primary development environment because it met all requirements: browser-based access, real file system, live preview of React+Vite builds, ability to pair with Claude for edits without losing visual feedback, and free tier. Staying in AI Studio would have constrained the project's ability to manage multiple test files and a backend structure. This decision enabled the scale-up from prototype to platform.
Kept handwritten research notes exactly as written — typos, uncertainties, casual observations stayed. Only structure and formatting changed, not content or voice. Cleaning up would have erased the actual thinking process. The uncertainty *is* the data. This decision became a model for documentation throughout the project: raw process is data.
Accepted comprehensive repository restructuring (consolidating templates, moving archived projects, creating missing ESF directories) intentionally. This was marked @default because the suggestion aligned perfectly with actual needs, but it's worth noting as intentional infrastructure work that supported M2+ requirements for documentation and process tracking.
Deliberately said "no" to healthcare provider access for the v1 platform. Adding provider authentication, access controls, and data privacy compliance would expand scope significantly when I needed to focus on getting the core loop working first: users take tests, see results, track progress. This protected my ability to finish within the quarter timeline by keeping scope focused on what matters *this quarter*. Provider integration became a Post-M3 feature for future iterations.
Sessions 7–12: Building three tests, creating standards, establishing visual language.
Rejected emoji icons (⚡✋🧠🤝👁️🎤) for test cards because they signal game-like experience rather than assessment tool. Research showed highly gamified designs (emoji, bright colors) risked feeling "too fun to feel trustworthy" for clinical contexts. Switched to clean SVG illustrations and intentional design system instead: dark blue + white + light blue for professional but approachable feel.
Rejected generic template aesthetic (basic colors, minimal personality, standard card layouts) and specified more opinionated design: dark blue + white + light blue color system, vector illustrations, split-card design with sophisticated visual hierarchy, micro-interactions. Template designs don't engage users. Building curiosity requires intentional design system.
Rejected standard hover states (shadow and lift) and specified animated illustrations showing test mechanics on hover. This serves both UX (clarity) and engagement (curiosity) — users can see what they're about to experience before committing. Hover should educate, not just provide visual feedback.
Removed all @keyframes animations from SVG illustrations and requested static, polished look first. Animation is a layer of polish that should come *after* structure is proven. Animations can be added intentionally later once the core experience is locked. This follows the principle: structure first, then movement.
Consciously chose vanilla HTML/Canvas for bilateral coordination test instead of continuing with React. React abstractions were getting in the way of direct control over canvas transformations, performance optimization, and debugging hand tracking logic. Vanilla canvas gave me full control over frame-by-frame game loops and immediate DOM manipulation for tracing bugs.
Went through 5 iterations of troubleshooting double-mirroring problem before discovering the root cause: CSS `scaleX(-1)` flipped the display, but canvas context transforms were flipping the image being drawn — two flips canceling each other. The solution: keep only CSS transform, remove canvas transforms, adjust coordinates with `(1 - indexTip.x) * width`. This taught the principle of understanding transformation pipelines: don't flip at multiple levels.
Abandoned counter-based dot tracking and rewrote spawning logic to count actual dots in the game array. Separate counters create single-source-of-truth problems — if the counter says "3 dots" but the array has 2, logic breaks. Always derive all state from actual game objects: `const count = enemies.filter(e => e.color === colorToCheck).length`. This prevents sync bugs.
Made color matching explicit: Left hand can ONLY grab red dots, Right hand can ONLY grab teal dots. This is explicit over implicit for debuggability and clarity. When bugs appear, it's immediately clear what's happening. The rule is explicit in the code.
Added an `isCollected` flag guard to prevent double-counting when both hands process the same dot in the same frame. Game loops are asynchronous — both hands could grab the same dot simultaneously. Defensive programming: if both hands process it, the second gets ignored because `isCollected` is already true. Ensures each dot counts exactly 1 point.
Requested comprehensive standardization of header layout, game flow, and UI patterns across all games with a 7-part specification document. This represented a fundamental shift: instead of designing each game's UI individually, every game now inherits a proven structure, freeing focus to just game-specific mechanics. Standardization is not loss of creativity — it's the opposite. By locking down patterns, I could build 5–10 different assessment tools instead of 2–3.
After Session 8, I forgot to commit prototype-5 to git. File state became corrupted with lock files and broken references. Rather than patch it incrementally, I explicitly chose to rebuild from zero. Lesson: **commit at end of every session, no exceptions.** The cost of rebuilding was high, but the lesson was foundational — version control is foundational to the entire quarter question. If you lose code frequently, you can't reliably build a platform that tracks anything.
Chose per-hand score tracking (`testLeftScore` / `testRightScore`) instead of just a total score. The clinical value *depends* on that breakdown — knowing left=8 and right=12 tells you about hand dominance and asymmetric patterns. Just knowing total=20 loses that insight. This is about measurement validity, not UX preference.
Specified `min(left, right) / max(left, right) * 100%` as the symmetry metric instead of a generic difference or percentage. This precise formula shows hand balance (100% = perfect, 0% = completely imbalanced), normalizes to 0–100%, is symmetric regardless of dominant hand, and is clinically meaningful. Precision in measurements matters.
Removed the bottom header/footer bar and made game full-screen. A test environment feels different from a homepage — removing visual distractions supports test validity (users focus on the task, not peripheral UI). Exit button + game title in top-left, HUD box in top-right. Everything else removed. Visual hierarchy supports test validity.
Accepted comprehensive bug fixes (restructuring `removeDot()` to separate state from visual feedback, fixing grab conflict logic, reducing grab range from 120 to 40 pixels) because all aligned with core principles already established about separation of concerns and explicit logic. These fixes were consistent with earlier learning about transformation pipelines and state synchronization.
Specified that left-hand dots spawn only on left side, right-hand dots only on right side. This removes a confound: when your hand blocks your view of targets, you're measuring hand positioning + eye-hand coordination, not pure bilateral coordination. Removing confounds makes measurement cleaner and more valid. In assessment design, isolate the specific skill being measured.
Rejected auto-generated results showing only totals. Demanded full-width comprehensive results showing detailed 60-second timeline (each grab as a colored dot), per-hand metrics (dots collected, grab rate, dominance), and coordination analysis (hand switches, consistency %, fatigue index). The same number of dots with perfect alternation = excellent coordination. Same number with erratic switching = poor coordination. Totals hide the pattern. Results screens should reveal what the test actually measures.
Changed test duration from 90 seconds to 60 seconds. Shorter tests are less fatiguing, meaning grab quality stays consistent throughout. A 60-second test measures raw bilateral coordination without introducing fatigue as a confounding variable. Test duration affects what you measure — shorter test isolates the skill you're testing. Keep tests focused and short. Measure one thing well, not multiple things poorly.
Explicitly kept experimental prototype-4 as a documented artifact instead of deleting it. Experiments are data, not waste. Prototype-4 contains learning: what visual direction didn't work (gradients, animations), why it didn't work (contradicted minimal aesthetic), and what the save-experiment-revert workflow looks like. Failed experiments have teaching value — deleting them would erase evidence of how you reached conclusions.
Kept placeholder cards visible with "Coming Soon" text instead of removing them. Removing cards hides the platform's roadmap. Keeping them visible shows users what's planned without pretending it's ready. This is a principle about transparency: be honest about what works and what doesn't. Hiding unfinished features erodes trust.
Corrected Claude to serve prototype-5 (working vanilla HTML version) instead of prototype-1-react. File ownership and clarity is non-negotiable — in a project with multiple prototypes in different formats, knowing which is canonical prevents wasted effort. Editing the wrong file means debugging wrong code and discovering problems too late.
When offered a menu of backend options (Firebase, Supabase, Node/Express), I rejected premature choice and asked instead "what IS a backend?" Taking time to understand the three-tier architecture (frontend / API layer / database) and why Supabase handles auth separately made every backend problem afterward understandable rather than magical. Knowledge precedes decision.
Didn't immediately move to Supabase implementation. Instead, deepened the conceptual understanding first. This established the mental model: the backend is invisible to users, doesn't change the UI at all, and sits underneath what I've already built. When I finally began setup, I understood *why* each step was necessary because I knew the architecture. Understanding compounds forward — it's what caught the Session 19 security issue.
Replaced email + sign out in header with small circular profile icon linking to dedicated history page. Account management shouldn't live in nav header — that's an information architecture principle. The nav should be minimal and focused on navigation. Account management (email, sign out, settings) is a different concern. Mixing them creates visual noise. Separate concerns in IA.
Replaced large SVG gesture icons (hundreds of lines of hard-to-read markup) with emoji equivalents. SVGs technically worked but added complexity, were inconsistent across browsers, and were hard to modify. Emoji are simpler, more consistent, universally supported. Tool selection principle: choose the simplest tool that works. Don't generate code when a built-in feature is simpler and more reliable.
When Rhythm test flow skipped practice phase, I explicitly fed back GAME_STANDARD phases and required the new flow to match: Welcome → Practice → Ready → Countdown → Test → Results. Consistency is a usability feature. Users learn the pattern once and apply it everywhere. Enforcing standards across all components ensures predictable experience. Standards aren't limiting — they guarantee every test is intuitive.
Sessions 13–16: Building and validating the testing platform.
Rejected "test with one hand only" as an edge case. Testing bilaterally with only one hand doesn't test bilateral coordination at all. Breaking tests must stress what you measure. Redirected to actual two-hand confounds: hands overlapping, one dominant, rapid pinching, hand fatigue. Edge cases must test the feature being measured, not arbitrary scenarios.
Replaced frequency-domain analysis (which returned zero taps for short taps because frequency analysis needs duration to build frequency content) with time-domain peak detection. Time-domain directly measures amplitude spikes — exactly what taps create. Frequency-domain approach was fundamentally mismatched to the problem. Fundamental algorithm shift that unblocked the entire game. When measuring transient events (taps, clicks), analyze time domain, not frequency.
Replaced Tailwind CDN + async Supabase scripts with inline CSS only. CDN loading is asynchronous and unreliable — if Tailwind loads after the page renders, styles never apply. Inline CSS guaranteed no network round-trips, consistent rendering, works offline. Shifted from modular architecture with external dependencies to self-contained deployment model. Single file = easier to share and test.
Refactored scattered functions defined at script scope into a single centralized `app` object containing all state and methods. Event handlers couldn't find functions; state ownership was unclear. Centralizing to one object guaranteed one source of truth. This is the state machine pattern — fundamental shift from scattered functions to centralized state machine. Guarantees predictable behavior.
When Rhythm Sync test had unique visual design different from Reaction Time and Bilateral Coordination, I extracted exact visual patterns from both existing tests and applied them: same color scheme, header layout, HUD box styling, animation keyframes, overlays, typography, spacing. Inconsistency requires users to relearn interface for each test. Consistency enables cognitive load reduction and supports building 5–10 tests without redesign.
When audio detection worked but users couldn't tell if input was being recorded, implemented three parallel feedback channels: large text feedback (Perfect/Good/Off/Miss), circle flash animation (tied to detection), real-time tap counter. Multiple modalities together create unambiguous feedback. No single channel is enough. Standard UX practice, worth recording as pattern for future tests.
Redesigned graph navigation from separate Previous/Next buttons to animated pill/toggle component. Single pill component with animated state transitions feels more cohesive, takes up less space, is more modern. Choose UI components that are both compact *and* immediately understandable. Don't default to traditional button layouts if a more unified component exists.
Made instructional text conditional based on context. "Click any point to see metrics" is useful on history page (users might not know what's clickable) but unnecessary on post-test results page (users are naturally exploring). UX instructions should adapt to user context, not apply universally. Respect for user context improves the same feature depending on where the user is in the flow.
Changed right hand label from "Teal" to "Blue" (the actual color is #1a3a8a blue, not teal). Inaccurate terminology creates confusion and breaks color consistency. Call things by their actual names. Users learn better when accuracy is consistent. Taxonomy precision supports credibility.
Changed "Fastest Response" color from green (#15803d) to blue (#1a3a8a — matching right hand color from bilateral test). Green + red mixing optically creates brownish tone at line overlaps, making visualization harder to read. Blue + red remain visually distinct even when overlapping. Using same blue creates visual coherence across platform. When designing multi-line graphs, test color combinations for visual clarity at intersections.
Sessions 17–19: Code review, fine motor precision, security hardening.
Rejected pagination controls and implemented simple scrollable container instead. For 5–20 results, scrolling feels natural and mirrors modern interfaces. Pagination requires users to decide which page and adds extra clicks. A scrollbar is intuitive and requires zero interaction to understand. Simple beats complex when both work.
Kept double-flip (camera feed AND coordinates both mirrored) to create intuitive mirror-world experience: "I move my hand right → dot moves right." Pure technical correctness would have opposite behavior but be confusing. User mental model > Technical purity. A system that's intuitively correct is better than one that's technically correct but confusing.
Removed all timers from fine motor test. Adding timer creates anxiety, which impairs fine motor function. Without time pressure, test reveals true motor control. Measure one thing well, not many things poorly. A test for precision shouldn't also measure speed. Keep tests focused and short.
Added pulsing glow on target portal and portal color change when reached. Clarity > Minimalism in a teaching context. Players need clear feedback about where they're going and whether they reached it. The goal is to help players understand the task. Pure minimalism assumes understanding without guidance.
Kept finger path in hand-space (raw coordinates, no offset) while keeping maze in maze-space (offset applied). Don't apply all offsets uniformly for consistency — apply transformations based on *what the element represents*. Different elements serve different purposes. Cursor needs to match path trail, so both use same coordinate system. Semantics before uniformity.
Made finger path opaque (0.85 opacity, 3.5px width) instead of subtle and semi-transparent. The path proves system is tracking correctly. If too faint, players lose confidence. A clear path provides real-time confirmation. In teaching context, feedback clarity > visual minimalism. The path is a teaching tool, not decoration.
Explicitly shelved Rhythm Synchronization as M5+ future work instead of counting it as M4 complete. Three polished user-tested games is more credible than four games where only three were tested. Quality threshold > Feature count. "Complete" means passing the same validation that other features pass. This protected the integrity of M4 findings.
Renamed "Finger Maze Test" to "Fine Motor Precision Test." Names shape user understanding — the interface is a maze, but the measurement is fine motor precision. For a "teaching" platform, naming clarity is essential. Parallel naming convention across tests (Reaction Time Test, Bilateral Coordination Test, Fine Motor Precision Test) communicates purpose. Name for meaning, not for mechanism.
Removed the entire HUD element (top-right instruction text and timer display). Instructions already appear on welcome screen — HUD is redundant and creates visual noise during motor control test. Remove before adding. If an element is redundant or distracting, eliminate it. Simplification through deletion.
Changed from "Fatigue Indicator: Your accuracy dropped 15%" (frames as deficit/failure) to "Movement Control: 78% excellent" (frames as capability/strength). Same data, different framing. In motor learning test, control framing is more useful and empowering. Frame feedback around capability, not deficit. Positive framing increases engagement and motivation. Represents shift from deficit to capability framing.
Session 20: Presenting work and completing the process book.
Rejected the AI-drafted artist statement scaffold entirely and wrote a complete statement from my own reflections. The initial draft was well-articulated but was Claude's voice, not mine. An artist statement is the capstone reflection of the artist's own thinking, not a place for AI scaffolding. I answered three guiding questions about the project and built the statement from those answers, emphasizing: the PRD as foundation, tracking change over time, maintaining autonomy while working with AI, understanding the backend, and learning to trust instincts. The artist statement needed to be accurate to my own voice and thinking, not just articulate.
These 54 Records of Resistance track the decision-making arc of the Baseline project across 6 milestones. Each record documents a moment where I deliberately pushed back on AI suggestions, refined a recommendation, or made a creative choice that differed from what was offered. Organized by milestone phase, they show recurring patterns: scope protection, measurement accuracy, explicit over implicit logic, understanding before building, design intentionality, standardization enabling scale, and choosing proportional solutions. The full compiled documentation is available in /records-of-resistance/COMPILED_RECORDS_OF_RESISTANCE.md.