Applied AI product builder working where operational complexity meets agentic systems.
I build AI systems for consequential workflows: workflow decomposition, tool contracts, structured outputs, evals, human approval, observability, and failure recovery.
My background is in commercial construction operations. Most of my production systems operate inside real business workflows and cannot be open-sourced, so this profile contains measurement tools, sanitized reconstructions, and public experiments around the same problems.
| System | Result |
|---|---|
| AI bid review | $221,915.70 surfaced across $3.44M reviewed · 34 jobs · 6.4% verified catch rate · zero escapes past human review |
| Quote import workflow | 30–60 hours/week returned to a 12-person estimating team · 12 releases · no rollbacks |
A local CLI for measuring Claude Code context pressure, delegation, model distribution, and token economics from real transcripts.
The accompanying field study examines what happened when fixed-contract specialist agents, context guards, and explicit session handoffs were introduced into a production development workflow.
A synthetic reconstruction of a real quote-to-order workflow, built to explore what production agent systems need beyond the demo: governed tools, human approval, evals in CI, audit trails, and observability.
Currently being built in public.
Local-first Windows dictation. Hold a key, speak, release, and text appears wherever your cursor is. On-device speech recognition, no account, no telemetry, and optional local-LLM enhancement.
Agent reliability Evals, regression testing, failure recovery, tool contracts, and human decision boundaries.
Agent economics Context pressure, orchestration, model routing, delegation, and return on inference.
Applied AI systems Turning messy human workflows into systems that can be measured, trusted, and improved.
Local-first intelligence Persistent agent systems where the user owns the state, tools, history, and operating environment.
workflow → decomposition → tools → state → agent → evals → human gate → outcome → feedback
The interesting question is not whether an AI system can complete a demo.
It is whether you can put it inside real work, understand when it fails, and make it better without breaking what already works.
More work and case studies at josiahbujanda.com

