Test Architect | AI Quality Engineer | Forward Deployed AI Engineer | Agentic AI | RAG & LLM Evaluation | MCP | Playwright | API Automation | CI/CD
I architect enterprise Quality Engineering and AI Assurance systems spanning traditional automation and production AI. My focus is not only whether a test passed, but whether there is sufficient evidence to prove that an AI system is reliable, grounded, safe, authorized, observable, governed and ready for release.
My portfolio connects:
Data → RAG → Prompt → Model → Agent/MCP → Security → Evaluation → Human Review → Release → Observability → Incident Response → Governance
Evidence first. AI assists engineering decisions; it does not replace engineering controls.
Enterprise AI Quality Portfolio & Recruiter Showcase — the recruiter front door to a 24-application AI Quality Engineering and Assurance ecosystem.
| System | What it demonstrates | Live demo |
|---|---|---|
| AI Release Assurance & Governance Control Plane | Policy-as-code gates, evidence bundles, human approval, rollback contracts and certification | Open |
| Agent Identity, MCP & Tool Governance Studio | Agent identity, MCP attestation, tool scopes, runtime authorization, approvals and violations | Open |
| RAG Evaluation Workbench | Retrieval traces, Precision/Recall/MRR/NDCG, groundedness, hallucination risk and regression | Open |
| AI Observability & Production Monitoring Center | AI traces, SLOs, drift, release correlation and production reliability | Open |
| AI Red Team & Safety Evaluation Center | Adversarial evaluation, safety regression, severity and remediation evidence | Open |
| FHIR AI Quality Lab | FHIR validation, SMART scopes, PHI handling, clinical provenance and grounded AI evaluation | Open |
flowchart LR
D[Data] --> R[RAG]
R --> P[Prompt]
P --> M[Model]
M --> A[Agent / MCP]
A --> S[Security]
S --> E[Evaluation]
E --> H[Human Review]
H --> G[Release Gate]
G --> O[Observability]
O --> I[Incident Response]
I --> V[Revalidation / Governance]
A production AI transaction should be explainable through an evidence chain such as:
Source → Chunk → Retrieval → Prompt → Model → Agent → MCP Server → Tool → Authorization → Evaluation → Safety → Human Approval → Release → Production Trace → Incident → RCA → Rollback → Revalidation
| Priority | Repository | Engineering evidence |
|---|---|---|
| 1 | Enterprise AI Quality Engineering Platform | Unified LLM/RAG/agent/MCP evaluation, datasets, security, observability and hard release gates |
| 2 | Playwright Enterprise Test Framework | TypeScript UI/API automation, cross-browser execution, accessibility, CI/CD and evidence-rich quality gates |
| 3 | Agentic Quality Engineering Platform | Governed QE agents, explicit state, bounded tools, RBAC, human approval and Playwright execution |
| 4 | RAG & LLM Evaluation Lab | Hybrid retrieval, reranking, groundedness, hallucination, citations, latency, tokens and regression evaluation |
| 5 | AI Agent Evaluation Framework | Task success, tool use, trajectories, grounding, safety, approvals and recovery evaluation |
| 6 | API & Integration Testing Framework | REST, GraphQL, Pact, RBAC, fault injection, events, retries and idempotency |
| AI Quality & Assurance | Agentic AI | Enterprise QE |
|---|---|---|
| LLM/RAG evaluation | Agent evaluation | Playwright + TypeScript |
| Groundedness & hallucination | MCP governance | API & contract testing |
| Golden datasets | Tool authorization | CI/CD quality gates |
| AI red teaming | Human-in-the-loop controls | Performance & reliability |
| Model/release risk | Runtime traces | Evidence & reporting |
| AI observability | Agent identity & delegation | Test architecture |
- Portfolio Showcase — understand the complete ecosystem.
- RAG Evaluation Workbench — inspect retrieval and grounding evidence.
- Agent Identity, MCP & Tool Governance — inspect agent/tool authorization.
- Release Assurance & Governance — convert evidence into a release decision.
- AI Observability — follow the system after deployment.
- FHIR AI Quality Lab — see domain-specific healthcare AI assurance.
The objective is not to show 24 disconnected dashboards. It is to demonstrate one production AI assurance story.
My strongest alignment is with roles involving:
- AI Quality Engineer / AI Quality Architect
- Test Architect / Quality Engineering Architect
- Forward Deployed AI Engineer
- Agentic AI Quality Engineer
- LLM / RAG Evaluation Engineer
- AI Assurance / AI Governance Engineering
- AI Test Automation Architect
- Deterministic before probabilistic — use schemas, contracts and exact evidence wherever possible.
- Hard gates remain hard — critical safety, authorization and correctness failures cannot be averaged away.
- Trace the whole system — data, retrieval, prompts, models, agents, tools and releases should be attributable.
- Human accountability — high-impact actions retain approval and audit evidence.
- Production feedback becomes regression evidence — incidents should strengthen future evaluation suites.
- Transparent portfolio claims — synthetic/reference demonstrations are kept distinct from measured enterprise delivery outcomes.
Agentic AI · MCP · RAG · LLM Evaluation · AI Observability · AI Governance · AI Safety · Playwright · TypeScript · Python · API Testing · CI/CD · Azure OpenAI · FHIR · Performance Engineering
- Live portfolio: AI Quality Engineering Portfolio
- LinkedIn: Ashok Kumar Manohar
- GitHub: ashokmanohar-ai
- Portfolio evidence: Verified Evidence Index
Engineering Quality for the AI Era.
Building quality systems that turn AI behaviour into reviewable engineering evidence.
