JORDAN TRUONG/ AI ENGINEER
SECTION01 / INDEX

GOODYEAR, AZ · REMOTE (US)
HYBRID PHOENIX

BUILDS AGENTS THAT SHIP.

Jordan Truong — AI Engineer in Goodyear, AZ. Open to remote-US or Phoenix-area roles.

Agentic engineering — formal specs, CI/CD gates, eval-gated reliability. Not vibe coding.

Production agentic AI: multi-agent systems, RAG pipelines, and the infrastructure beneath them. Solo-built and run in daily production, end to end, from a terminal in Arizona.

The arc: twelve years a salon owner-operator. Then roughly four years in enterprise IT. Now AI engineering, shipping software that runs itself. Same instinct throughout: make the hard thing dependable.

The Résumé

SECTION02 / RÉSUMÉ20-SECOND SCAN · EVERY LINE PROVABLE

Core

Multi-agent orchestrationContext engineeringMCPRAG / GraphRAGLLM evalsGuardrails

Stack

PythonTypeScript / JavaScriptReactNext.jsNode.jsSQLPostgreSQLDocker

Domain

HIPAAPCI DSSNIST 800-53 / CSFLangfuseAWS BedrockGCP Vertex AI
2025 → PRESENTFounderBlu Print Solutions, LLClive demos
AUG 2024 → AUG 2026Customer Service Representative, Multi-CampaignValor Globalwalkthrough
FEB 2023 → APR 2024Service Desk Support AnalystMassage Envy (Corporate)detail
MAY 2022 → JAN 2023IT Support SpecialistAlorica (HMH + Dell campaigns)detail
OCT 2010 → MAR 2022Independent Business OwnerSelf-Employedthe arc

Also live and provable, found in tonight's audit: /opsLangfuse-traced cost dashboard with per-request pricing · Anthropic Academy ×15, Claude Partner — Claude Code, and Google AI Professional Certificate.

Every role above links to a live demo or its full walkthrough — never a bare claim.Same content as a 2-page PDF ↓

Experience

SECTION03 / TIMELINE
2025 → NOW
Founder & AI EngineerBlu Print Solutions

Solo-built multi-agent VPS, SoloInvoice SaaS, RCM knowledge assistant. All in production.

MCPRAGGraphRAGMulti-agentSaaS
2024 → AUG 2026
Customer Service Representative — Multi-CampaignValor Global

Built production AI tooling across two enterprise campaigns: HIPAA RCM copilot (every scenario audited, 72 KB-grounded scripts, 27 RPM playbooks, 41 canned-note codes) and ACES Aid (162 flows, floor-wide adoption, 1,500+ indexed keywords). Both in daily floor use.

HIPAAAgentic flowsEvalsRCM
2023 → 2024
Service Desk Support AnalystMassage Envy (Corporate)

Managed provisioning across 1,000+ franchise locations. Led network/security audits with firmware rollouts sequenced across timezone waves.

ITServiceNowInfrastructure
2022 → 2023
IT Support Specialist — Multi-CampaignAlorica

Two enterprise campaigns: HMH education software (app/account triage, escalation) and Dell Technologies (hardware diagnostics via Zendesk, live Teams collaboration with senior techs, warranty logistics).

IT SupportZendeskSLA
2010 → 2022
Owner-OperatorSelf-Employed

Twelve years running a service business end-to-end: hiring, payroll, scheduling, P&L. Operational systems thinking at its most direct.

OperationsP&LPeople management

The Work

SECTION04 / RECEIPTS
01FLAGSHIP · PYPI · MIT · OSS

all41n14lla

Open-source MCP memory server. One command to install: pip install all41n14lla. A 90-test suite runs green in CI across Python 3.11 through 3.14.

  • MCP server
  • 90 tests · CI green
  • Py 3.11–3.14
  • MIT license
Python · MCP · PyPIOpen repo
02● 10 · PRODUCTION

The Builds

Ten agentic systems in daily production: healthcare + telecom copilots (every scenario audited, 162 flows), five in-browser demos — RAG, SQL, email triage, agent orchestration, semantic layer — SaaS with offline licensing, a 7-agent outreach engine, and the platform wiring all of it. Tap to browse.

  • Healthcare + Telecom AI · 2
  • In-browser agents · 5
  • SaaS + ed25519
  • Agency Engine
Python · Claude · Next.js · GroqBrowse all
03HARNESS / INFRASTRUCTURE

The 90% — Harness Architecture

Fifty-plus skills loaded on demand, fifteen MCP servers, auto-memory across sessions, and layered guardrails. The model is 10% of the system — the harness is the other 90%.

  • 50+ skills on demand
  • 15 MCP servers
  • Layered guardrails
  • Cross-session memory
Instructions · Memory · GuardrailsRead
04WIP · SALON NICHE

THE AGENCY ENGINE

Scout finds. Diagnoser maps. Builder ships. Filmer cuts. Checker validates. Pitcher sends. No staff. No retainer. $0.10/lead.

PROBLEM
Salons have no web presence and no time to fix it. Manual agencies charge $5K+ per client for work that's 80% repetitive.
SOLUTION
7 agents chain: Scout → Diagnose → Build → Film → Check → Pitch. Human touchpoint: payment only.
WHY BEST
Volume × personalization at scale. Only chained agents can run hundreds of targeted, compliant campaigns under $1/lead.
  • Scout → Film → Pitch
  • $0.10/lead · zero staff
  • CAN-SPAM compliant
  • ~$10/mo ops cost
Python · 7 Agents · LLM router · VercelBUILDING

Try It Live

SECTION05 / DEMOS

Or open the chat bubble (bottom-right) to ask me anything.

Every demo ships through typecheck, eval, and prompt-regression gates on every commit — and the guardrails are visible live, not just claimed: RAG declines instead of guessing, the Critic rejects unsupported claims, the semantic layer refuses off-catalog metrics.

LORA FINE-TUNE · $0 · ON-DEVICE

One experiment that didn't beat prompting — published anyway.

Fine-tuned Qwen2.5-0.5B locally on 8 examples from this site's own content, testing whether it could drop the ~190-line system prompt the live chatbot runs on. It couldn't. Real transcript from the eval run:

Q: "What is all41n14lla?"

Base model + system prompt — correct: "An open-source MCP memory server for AI agents, maintained by Jordan Truong, published on PyPI and GitHub."

Fine-tuned, no prompt — hallucinated: "...a web application for managing all 41 million users of the Netflix app..."

DEEP DIVE — 40 MIN

AI hosts break down my work through the lens of Google's New SDLC.

LOADING...

The Arc

SECTION06 / STORY
2014THE FLOOR

Twelve years on the floor.

Owner-operator of a salon and beauty business. Scheduling, payroll, inventory, hiring, client experience — all of it. The job taught systems thinking before I had a name for it: every constraint is a design problem, every bottleneck is a process failure waiting to be solved.

2022THE SWITCH

Pivot to enterprise IT.

Joined the workforce-management tech stack at scale. Claims processing, healthcare RCM, multi-team ops coordination. Learned what large-system reliability actually means when a misconfigured rule costs someone a reimbursement. Then LLMs became capable enough to matter.

2026● NOW — THE BUILD

Building AI that runs itself.

Shipped a PyPI MCP server, a live RAG+GraphRAG knowledge assistant, a multi-agent VPS brain, and a licensed SaaS product — all solo, all in production. The plan is to keep shipping until the systems are earning while I sleep.

Stack

SECTION07 / TOOLS

AI & Agents

  • Multi-agent orchestration
  • RAG / GraphRAG
  • MCP (Model Context Protocol)
  • Prompt engineering
  • Eval frameworks
  • Guardrails & abstention
  • Prompt-injection defense
  • Cost routing

Models & Cloud

  • Claude (Anthropic)
  • Gemini (Google)
  • GPT-4o (OpenAI)
  • Groq (Llama 3)
  • Perplexity (search)
  • Google Cloud / Vertex
  • Vercel Edge
  • VPS (Debian + systemd)

Build

  • Python 3.11–3.14
  • TypeScript / Node.js
  • React 19 + Vite
  • Next.js 15
  • FastAPI
  • PostgreSQL / Supabase
  • HelixDB (graph)
  • Stripe / webhooks

Certs & Security

  • Anthropic Academy (17 certs)
  • Google AI Professional Cert
  • CompTIA Security+ (active)
  • CompTIA A+ (active)
  • HIPAA compliance
  • NIST 800-53 controls
  • ed25519 offline licensing
  • Zero-PHI architecture

What Actually Runs This Site

SECTION08 / SYSTEM MAP

One diagram, drawn from the code the same day this was written — not a slide. Two entry points, one shared trust boundary, three honesty tiers on every demo answer.

VISITORjordantruong.comCHAT WIDGETapi/chat.jsIts own 6-layer defense1. keyword screen2. canary token3. fingerprint check4. anti-extraction5. online scoring6. adversarial taggingNOT abuseGuard — a purpose-built system for the one always-on routeGroqLlama 3.3 70B/demos.html5 demosapi/demos/{rag,sql,email,agents,semantic}.jsabuseGuard on all 5live · precomputed · client-retrievalSHARED TRUST BOUNDARYSupabaseRAG vectors + demo rate limitsLangfusetraces every model call, both paths
VERIFIED IN CODE, SAME SESSION

What this system is NOT: no PHI, no real user accounts, no payment data. A HIPAA-safe tool built at Valor is a separate system — not this one.

The Calls

SECTION09 / TRADEOFFS

Five decisions from the systems above, with the cost attached. Anyone can list what they shipped — these are the choices that had a price, including the ones that cost me something.

01

The model picks the metric. It never writes the query.

CONSTRAINT
Natural-language questions over real business data. A confidently wrong number is worse than no answer, and every answer has to be explainable afterward.
OPTIONS
Let the model write SQL and validate it · let it write SQL against a restricted view · publish a fixed metric catalog and let the model only choose from it.
CHOSE
The catalog. A validator can only reject SQL it anticipated; a generator that cannot produce arbitrary SQL has no unanticipated output to reject. The model is reduced to a classification problem and a deterministic compiler emits the query.
COST
Open-ended questions are genuinely gone — anything outside the catalog cannot be asked at all. I took a real product limitation over ever being confidently wrong. The catalog is also a maintenance surface: every new question is an entry, not a prompt tweak.
REVISIT WHEN
Catalog coverage becomes the top complaint. The honest next move is the restricted view behind row-level security — not loosening this one.
IN CODE
src/demos/lib/semantic-layer.ts · api/demos/semantic.js
02

I fine-tuned a model, it lost to prompting, and I published that.

CONSTRAINT
The chatbot needed tighter persona adherence. Whether that is a training problem or a prompting problem was genuinely unknown — which is exactly when people guess.
OPTIONS
LoRA fine-tune a local model · better prompting plus retrieval · move to a larger hosted model and pay for it.
CHOSE
Measure first. I built the dataset, ran the training, and evaluated before and after against the prompting baseline. Prompting won, so the fine-tune did not ship.
COST
Real time spent on a model that never shipped. I traded it for a measured number instead of an opinion. Shipping the fine-tune because I had already paid for it would have been the more expensive mistake.
REVISIT WHEN
When volume makes per-token cost dominate quality. A small tuned model competing on economics is a different question than one competing on quality, and I only answered the second.
IN CODE
fine-tune/train_lora.py · eval_before_after.py · eval_results.json
03

Fail open on structure. Never fail open on content.

CONSTRAINT
A Researcher → Writer → Critic pipeline where the Critic fact-checks the draft. It must always terminate, and it must not pass an unsupported claim. Those pull in opposite directions.
OPTIONS
Retry until the response parses · hard-fail the run on any malformed response · split the two and treat them differently.
CHOSE
Split them. A malformed envelope is accepted and the run continues; an unsupported claim forces a revision round, capped at two. Strictness about JSON shape buys nothing. Strictness about claims buys everything.
COST
The cap is a real hole — a claim that survives two rounds ships. I took a bounded, known error rate over an unbounded loop, because "sometimes it never finishes" is the worse failure in something a stranger clicks once.
REVISIT WHEN
When revision rounds start hitting the cap regularly. That means the Critic is fighting the Writer’s prompt, and the fix is upstream — not a higher cap.
IN CODE
src/demos/lib/orchestrator.ts
04

Every demo degrades honestly, and says which tier answered.

CONSTRAINT
A demo is opened once, by a stranger, on a phone, with no API key, and gets exactly one chance. It must never show an error screen — and must never quietly fake a live result.
OPTIONS
Require keys and show most visitors nothing · ship canned responses styled as live · run three real tiers and label which one answered.
CHOSE
Three real tiers with a source badge on every answer: live, precomputed, client-retrieval. The SQL demo runs genuine SQLite in WebAssembly; retrieval runs real client-side search over a bundled corpus.
COST
Three code paths and three sets of tests instead of one, and I tell visitors when they are seeing the cheap tier — strictly worse marketing, strictly better engineering.
REVISIT WHEN
Never, for a portfolio. In a paid product the tiers stay and the badge becomes an internal signal instead of user-facing chrome.
IN CODE
src/demos/lib/types.ts · retriever.ts · sqljs-loader.ts
05

A public model endpoint is an untrusted input surface. Including mine.

CONSTRAINT
A chatbot on a personal domain, open to anyone, wired to a paid model API — simultaneously a prompt-injection target, a cost-exhaustion target, and a data-extraction target.
OPTIONS
Rate-limit at the edge and trust the prompt · layered defense at the application boundary with an abuse guard on every route that reaches a model.
CHOSE
The layered defense — and the interesting part is the miss. A review pass found the guard wired into some model routes and not into three others, including the live retrieval route already in production. The design was right; the wiring was incomplete.
COST
Latency on every request and false positives on legitimately odd questions. More honestly: a layered design created a false sense of coverage, because "we have a defense" is not the claim "it is attached to every path."
REVISIT WHEN
Every time a new route reaches a model. That is the trigger — not a calendar. The guard is now a route-level requirement rather than a component you remember to call.
IN CODE
api/chat.js · api/rag-search.js · api/demos/*.js
SECTION10 / CONTACT

LET'S BUILD SOMETHING THAT SHIPS.

JORDAN TRUONG/ AI ENGINEERAll systems. All solo. All in production.
© 2026 Jordan Truong