Contract fit · AI Engineer · Claude Code · Agentic AI

I already work in the engineering model this role is trying to scale.

Claude Code, coding agents, LLM applications, RAG, APIs and workflow automation only become valuable when the surrounding engineering system is reliable. That layer — context, tools, tests, evals, controls, observability and production delivery — is where my strongest fit sits.

The short answer

AI coding is useful. AI engineering is the job.

I use coding agents as part of a production delivery system: understand the repo, frame the task, plan the change, implement it, run tests, inspect the diff, evaluate the result, document the decision and release through controlled automation.

The same pattern extends into agentic applications: models get bounded tools, retrieval gets measured, APIs remain deterministic, and security, approvals, observability and fallback paths are designed before scale.

Claude Code + Codex/CLI
60+ reusable AI coding skills
115-role agent workspace
Python + FastAPI
TypeScript + Next.js
RAG + vector retrieval
MCP + tool calling
Evals + human review
Job-spec mind map

AI Engineer — the whole brief in one compact view.

Hover or focus a branch for my evidence. The visible keywords mirror the role specification.

Very strong

AI coding

GitHub Copilot · Claude Code · code generation · intelligent SDLC

Mat evidence

Daily Claude Code + Codex/CLI, 60+ reusable skills, MCP, subagents, tests and release automation. Copilot-specific depth is transferable rather than overstated.

Very strong

LLMs + GenAI

OpenAI GPT · Azure OpenAI · Anthropic Claude · prompt engineering

Mat evidence

Production Claude, OpenAI and Azure OpenAI work with routing, retries, fallbacks, prompt optimisation and cost/latency controls.

Very strong

Agentic AI

Agents · copilots · workflows · LangGraph · CrewAI · AutoGen · Semantic Kernel

Mat evidence

AIOS, AIBO and ABC: planner/router/executor patterns, MCP tools, memory/state, typed hand-offs, approvals and recoverable fallbacks.

Very strong

Software engineering

Python · JavaScript/TypeScript · REST · microservices · scalable architecture

Mat evidence

Python/FastAPI plus TypeScript/Node/Next.js, REST APIs and production application architecture across enterprise systems.

Very strong

RAG + vector search

RAG · embeddings · semantic search · Pinecone · ChromaDB · Weaviate · Azure AI Search

Mat evidence

Embeddings, hybrid search, GraphRAG, LlamaIndex, pgvector, Pinecone, Weaviate, FAISS and Chroma with provenance and permission-aware retrieval.

Strong

Evals + governance

Evaluation · observability · monitoring · AI governance · security · responsible AI

Mat evidence

Tests, evals, tracing, auditability, GDPR/redaction, approval boundaries, provider fallbacks and human-in-the-loop controls.

Strong

Cloud + DevOps

Azure/AWS/GCP · GitHub Actions · Jenkins · Docker · Kubernetes · Terraform

Mat evidence

AWS, GCP and Azure exposure with GitHub Actions, Jenkins, Docker, NGINX and CI/CD. Kubernetes/Terraform depth should be validated rather than inflated.

Very strong

Enterprise + leadership

Enterprise integration · stakeholders · mentor developers · accelerate adoption

Mat evidence

15+ years across Allianz, RWS, Sainsbury's and D&B/Cogniflare bridging architecture, engineering, governance, stakeholders and delivery.

Target role

AI Engineer

Build AI-powered software with Claude Code, GitHub Copilot and modern LLMs — safely, measurably and at enterprise scale.

Build fasterAgents + RAGEnterprise APIsTest + evaluateGovern + observeShip + mentor
Requirement → evidence → move

The brief, mapped to production evidence.

The useful question is not whether a keyword appears on a CV. It is whether there is evidence behind it and a clear way to apply that evidence to the job.

Claude Code + AI-assisted SDLC
Very strong
Evidence

Claude Code is part of my daily engineering model: repo-aware analysis, planning, implementation, refactoring, testing, documentation and release workflows, supported by 60+ reusable skills, MCP integrations, subagents and hooks.

What I would do

Turn coding agents into a repeatable engineering system with explicit context, review, testing, permissions and release controls rather than isolated prompts.

GitHub Copilot
Strong adjacent fit
Evidence

My deepest day-to-day coding-agent experience is Claude Code and Codex/CLI. I have also delivered Microsoft Copilot and Copilot Studio solutions in enterprise environments, so the workflow and governance patterns transfer directly.

What I would do

Apply the same controlled AI-assisted delivery loop across GitHub Copilot while being precise about tool-specific experience rather than inflating it.

LLMs + GenAI applications
Very strong
Evidence

Production work with Anthropic Claude, OpenAI, Azure OpenAI and multi-provider model routing across live agentic platforms, including fallbacks, retries, cost/latency trade-offs and operational controls.

What I would do

Choose the right model boundary per workflow and keep providers replaceable behind tested service contracts.

Agents + workflow automation
Very strong
Evidence

AIOS, AIBO and ABC use multi-agent orchestration, MCP/tool contracts, router/planner/executor patterns, memory/state, typed hand-offs, human approvals and recoverable fallbacks across a 115-role workspace.

What I would do

Build bounded agents that can act through real APIs and business workflows while preserving auditability and human control where risk requires it.

RAG + vector databases
Very strong
Evidence

Hands-on retrieval work spanning embeddings, hybrid search, GraphRAG, LlamaIndex, pgvector, Pinecone, Weaviate, FAISS, Chroma and MongoDB-backed knowledge systems.

What I would do

Treat retrieval quality as an engineering problem: ingest, chunk, embed, retrieve, rerank, permission-filter, cite and evaluate against real user tasks.

Python / FastAPI + TypeScript
Very strong
Evidence

Hands-on backend and product engineering with Python, FastAPI, REST/JSON APIs, TypeScript, Node.js, Next.js and React, connecting models to production services and enterprise workflows.

What I would do

Ship thin end-to-end slices through the real stack so AI capability is tested as part of the application, not in a notebook beside it.

Evals + observability + governance
Strong
Evidence

Current production patterns include automated tests, evaluation, model routing, tracing, retries, caching, provider fallbacks, audit trails, incident-aware controls, GDPR/redaction and human-in-the-loop review.

What I would do

Measure task success, groundedness, regressions, latency and cost, then expose failure clearly enough that engineering teams can operate the system with confidence.

Enterprise delivery + technical leadership
Very strong
Evidence

15+ years delivering production software across Allianz, RWS, Sainsbury's, Dun & Bradstreet/Cogniflare and other enterprise environments, bridging architecture, engineering, governance and stakeholder delivery.

What I would do

Accelerate the team without creating dependency: establish patterns, pair with engineers, document decisions and leave reusable components and guardrails behind.

Production proof

Three examples that matter for this role.

Agentic architecture

AIOS

A live multi-agent platform with 115 defined roles, per-agent models/tools/memory, permissions, evaluations and provider fallbacks. The useful proof is not the number of agents; it is the operating model around them.

AI-assisted engineering

ABC

An engineering cockpit centred on 60+ reusable Claude Code skills covering build, test, deploy, analyse, ingest, sync and publish workflows — turning coding agents into repeatable delivery infrastructure.

Regulated enterprise AI

Allianz

Enterprise AI delivery spanning Copilot Studio assistants, voice/contact-centre automation, SharePoint/OneDrive grounding, Azure OpenAI and governance requirements including GDPR, redaction and auditability.

Engineering instinct

Let the model be probabilistic. Keep the engineering deterministic.

Agent behaviour can be flexible. Identity, permissions, schemas, API contracts, tests, deployment, logging, cost ceilings and rollback paths should not be.

AI-assisted SDLC

Repo context, bounded tasks, tests, diff review, CI/CD and reusable agent skills rather than one-off code generation.

Grounded applications

Retrieval quality, permissions, provenance, citations and evaluation before prompt theatre.

Agent control

Typed tools, explicit state, retries, approval gates and recoverable execution paths.

Enterprise operation

Security, auditability, monitoring, failure evidence, cost/latency budgets and an owned runbook.

First 30 days

One production engineering loop, then scale it.

The fastest route to trust is one real workflow with measurable quality and a team that understands how to operate it.

  1. 01

    Map the engineering loop

    Identify the highest-value developer and application workflows, current Claude Code/Copilot usage, codebase boundaries, model providers, data access and measurable acceptance criteria.

  2. 02

    Ship one vertical slice

    Take one real use case from repo context → agent/tool action → code/API change → automated tests → review → deployment, with observability and approvals included.

  3. 03

    Prove quality and safety

    Add evaluation, regression tests, groundedness checks, permissions, secrets boundaries, audit evidence, failure recovery and cost/latency measures around the workflow.

  4. 04

    Scale the pattern

    Turn the first successful loop into reusable skills, prompts, tools, templates, CI checks and team guidance so adoption compounds instead of fragmenting.

Questions I would ask

Six questions that reveal the real job.

  1. 01

    Is the primary outcome developer productivity, customer-facing AI applications, or both?

  2. 02

    How far has the Claude Code / GitHub Copilot rollout progressed today — individual usage, team standards, or platform-level enablement?

  3. 03

    Which use cases would make the first 90 days an obvious success?

  4. 04

    What are the main application stacks, cloud platforms and model providers already approved?

  5. 05

    Where must human approval remain mandatory, and where is bounded agent autonomy acceptable?

  6. 06

    How are AI-generated code quality, security, test coverage and production incidents measured today?

Bottom line

The fit is not one AI tool. It is the full engineering system around it.

I bring the architecture, hands-on coding, agents, retrieval, APIs, evals, controls and enterprise delivery needed to turn Claude Code, Copilot and modern LLMs into a reliable production capability the wider engineering team can own.

AI Engineer 100 · keyword index

100 concepts for the role, in one grid.

The job-spec terms plus the production concepts needed to understand them end-to-end. Hover or keyboard-focus any tile for its explanation.

Learn + test all 100
GitHub Copilot
Claude Code
LLMs
Generative AI
OpenAI GPT
Azure OpenAI
Anthropic Claude
Claude APIs
Prompt Engineering
Prompt Optimisation
AI Agents
Agentic AI
Copilots
Workflow Automation
LangChain
LangGraph
CrewAI
AutoGen
Semantic Kernel
Code Generation
AI-assisted SDLC
Developer Productivity
Python
FastAPI
JavaScript
TypeScript
C#
Java
REST APIs
Microservices
Scalable Architecture
Cloud-native
Enterprise Integration
RAG
Embeddings
Semantic Search
Vector Database
Pinecone
ChromaDB
Weaviate
Azure AI Search
Knowledge Repositories
Model Evaluation
LLM Evals
Observability
Model Monitoring
AI Governance
Responsible AI
Security
Human-in-the-loop
AWS
Azure
GCP
GitHub Actions
Azure DevOps
Jenkins
Docker
Kubernetes
Terraform
CI/CD
APIs
Mentoring
Stakeholder Use Cases
Enterprise AI
MCP
Tool Calling
Function Calling
Multi-agent Systems
Router / Planner / Executor
Memory + State
Subagents
Skills
Hooks
Guardrails
Prompt Injection
PII Redaction
Secrets Management
Least Privilege
Audit Trail
Groundedness
Hallucination
Golden Dataset
Regression Testing
Tracing
Latency
Token Cost
Model Routing
Provider Fallbacks
Retries
Caching
Idempotency
API Contracts
Schema Validation
OpenTelemetry
Git
Pull Requests
Unit Testing
Integration Testing
Architecture Decision Records
Technical Leadership
Roll the dice