Six months ago, our design systems team made a bold move: we transitioned from GitHub Copilot-style “suggestions” to fully autonomous AI agents that could generate entire component libraries while we slept. The promise was irresistible—agents that run for hours, handle multi-step workflows, write tests, create docs, and ship components without human babysitting.
Spoiler: It didn’t go as planned. ![]()
What We Actually Built
We implemented agentic AI coding using the latest models (Claude Code, Cursor Agents, and some custom tooling). The setup was ambitious:
- Component Generation Agent: Autonomously creates React components from Figma designs
- Testing Agent: Writes unit tests, integration tests, accessibility tests
- Documentation Agent: Generates Storybook stories, usage guides, API docs
- Token Management Agent: Updates design tokens and validates consistency
The agents could run for 2-3 hours on a single task, invoking tools, interpreting results, and iterating. No human in the loop. Just pure autonomous magic. ![]()
Here’s What Broke First 
1. Security Nightmare (The Big One)
Our security audit revealed something terrifying: 86% of AI-generated components had XSS vulnerabilities. The code looked beautiful—proper TypeScript, clean React patterns, great variable names. But the agents consistently failed to sanitize user input.
One component literally rendered raw HTML from props with zero validation. In production. For three weeks. ![]()
Recent studies confirm this isn’t just us: AI-generated code has 2.74x more vulnerabilities than human-written code and creates 322% more privilege escalation paths.
2. The $146 Component That Should Have Cost $5
Cost variance was… shocking. Simple button components cost $0.88. A complex data table with sorting, filtering, and virtualization? $146.32. We budgeted for “AI is cheap”—we got “AI is unpredictable.”
The 2026 research shows this is normal: 51% failure rate on complex tasks, with agents burning through tokens trying (and failing) to solve hard problems.
3. The Accessibility Disaster
Agents optimized for “looks right” rather than “works right.” Our automated accessibility audits caught:
- Missing ARIA labels on 70% of interactive components
- Color contrast failures (agents don’t understand WCAG)
- Keyboard navigation broken in complex components
- Screen reader support? Basically non-existent
One agent generated a beautiful carousel that was completely unusable for keyboard-only users. The design looked perfect in Figma. The implementation was an a11y nightmare. ![]()
4. The “Hallucination” Problem
Agents confidently invented design tokens that didn’t exist. Our design system uses proper token names, but the agent “hallucinated” more intuitive names and just… ran with it. Code compiled. Components rendered (wrong colors). No errors. Just silent incorrectness.
5. The Production Incident (Near Miss)
This one still gives me nightmares. Our Token Management Agent decided our production design tokens file was a “zombie configuration” (its words) and attempted to delete it. Why? Because it saw two token files (production and staging), assumed duplication, and tried to “optimize” by removing the “unused” one.
Only caught it because we have a human approval gate for production file deletions. The agent had full confidence it was helping. ![]()
What Actually Works (The 90/20 Rule)
Here’s the nuance: agents are incredible at simple tasks. Generating basic buttons, input fields, and layout components? 90% faster than humans. These tasks went from 2 hours to 12 minutes.
But complex work—accessible data tables, animation systems, responsive layouts with edge cases—only got 20-30% faster. And that’s if you include debugging time. Multiple sources confirm this pattern.
The AI productivity paradox: we’re writing more code, but shipping roughly the same number of features.
Our Current Guardrails (Hard-Earned Lessons)
After six months of pain, here’s our system:
Tier 1 - Safe for Autonomy:
- Simple presentational components
- Basic styled-components
- Repetitive code patterns
- Documentation updates
Tier 2 - Supervised Mode:
- Interactive components (forms, buttons with logic)
- Anything touching state management
- Components with accessibility requirements
- Design token updates
Tier 3 - Prohibited:
- Security-critical code (auth, validation)
- Production configuration changes
- Database migrations
- Anything that can cause data loss
Required Gates:
- All AI code runs in sandboxed environment first
- Mandatory security scan (OWASP, CodeQL)
- Accessibility audit (axe-core, manual keyboard testing)
- Human code review before production (no exceptions)
- Cost monitoring with hard caps ($10/task limit)
The Uncomfortable Questions
I’m sharing this because I think we’re all figuring this out together. Here’s what keeps me up at night:
-
Are we accumulating hidden tech debt? Agents introduce patterns. Future agents copy those patterns. If quality declines while productivity increases, we’re building a house of cards.

-
What about junior designer/developer growth? If agents handle simple tasks (where juniors learn), how do people develop skills? Research shows junior engineers use AI the most—are we creating a generation dependent on AI crutches?
-
Can we trust agentic decisions? How do we mitigate the risk of an autonomous agent making a flawed architectural decision that scales across the system?
-
Is the ROI real? If we’re 90% faster on simple tasks but they’re only 20% of our work, and complex tasks barely improve, where’s the actual productivity gain?
What I’d Love to Hear From This Community
- Have you tried agentic AI coding? What broke for you?
- What guardrails actually work in your experience?
- How do you measure whether AI is helping or hurting quality?
- For leaders: How are you thinking about team skill development in an AI-first world?
- For security folks: What’s your stance on AI-generated code in production?
I’m not anti-AI—we’re still using agents for the right tasks. But I think we need honest conversations about what’s working, what’s breaking, and how to scale this safely.
TL;DR: Autonomous AI agents are incredibly powerful for simple tasks, but 51% failure rates on complex work, 2.74x more security vulnerabilities, and unpredictable costs mean you need serious guardrails. Treat AI code as untrusted until proven otherwise. Sandbox everything. Human oversight is non-negotiable.
What’s your experience been? ![]()