← All case studies
Case study

Multi-agent system for data operations and support

Agents that query, diagnose and re-run integration jobs in production.

// Role
Integrations Engineer, customer-facing technical lead
// Context
Fintech company in financial data reconciliation and automation (private project)
// Stack
Google ADK, LangGraph, Gemini, Text-to-SQL, DynamoDB, PostgreSQL via MCP, AWS EKS, AWS Batch, AWS CloudWatch, ticketing system
#

Problem

Every customer came with different data sources, and every integration created recurring operational work: questions about job status, pipeline failures to diagnose by reading logs, and re-execution tickets that someone had to validate, launch and close by hand.

As an Integrations Engineer I was customer-facing. I listened to the problem in the meeting, understood the customer's source system and built the connection. After delivery, the cost was in daily operations: the same kind of ticket, several times a day, with the same diagnostic pattern.

The design question was: which part of that loop (query, diagnose, re-execute) can agents automate without losing control over customer data.

#

Architecture

A Router + Specialists pattern on a stateful graph (LangGraph, Google ADK):

flowchart TD
User["User Events (Slack / Zendesk / CLI)"] --> Router

subgraph EKS["Agent Services on AWS EKS"]
    Router["AgentRouter (gemini-2.0-flash)"]
    Router -->|"DATA_QUERY"| SQL["JobsSQLAgent"]
    Router -->|"DOCUMENTATION"| Docs["DocsRAGAgent"]
    Router -->|"DIAGNOSTIC"| Diag["SupportDiagAgent"]
    Router -->|"REEXECUTION"| Reexec["ReexecutionAgent"]
end

SQL --> DB[("PostgreSQL via MCP")]
Docs --> KB["Markdown Knowledge Base"]
Diag --> DB
Diag --> CW[("AWS CloudWatch Logs")]
EKS -.-> DDB[("DynamoDB (State & Memory)")]

Reexec --> API[("Backend API")]
API --> Batch[("AWS Batch")]

classDef default fill:#0f172a,stroke:#38bdf8,stroke-width:1px,color:#e2e8f0
classDef router fill:#881337,stroke:#f43f5e,stroke-width:1px,color:#fecdd3
classDef specialist fill:#1e1b4b,stroke:#818cf8,stroke-width:1px,color:#e0e7ff
classDef storage fill:#022c22,stroke:#10b981,stroke-width:1px,color:#d1fae5

class Router router
class SQL,Docs,Diag,Reexec specialist
class DB,CW,API,Batch,KB,DDB storage
Routeran LLM classifies the intent of the incoming event (ticket, chat, CLI) and delegates to the right agent. All agents share the same session context.
Query agent (Text-to-SQL)answers operational questions about jobs by querying PostgreSQL through an isolated MCP toolbox. It validates the schema and runs read-only, with curated query sets.
Documentation agent (RAG)answers from the Markdown technical knowledge base. It uses full-document analysis instead of vector search when exact requirements must be extracted.
Troubleshooting agentdiagnoses failures by combining CloudWatch execution logs with relational metadata.
Re-execution agent (orchestrator)reads the ticket, validates requirements, relies on the query and troubleshooting agents, and triggers the re-run through the backend API. The job runs on AWS Batch - the agent never touches AWS directly.
State and memorysessions in a NoSQL store with TTL; semantic memory and execution traces in DynamoDB.
Key Architecture Decision

Key decision: the re-execution agent has no query or diagnostic logic of its own. It calls the specialists, so each piece can be evaluated separately.

#

Results

Measured with my own evaluation, with no customer data exposed:

Query agent
95%correct responses, evaluated manually through the adk run web interface
Troubleshooting agent
98%+accuracy in my evaluation
Re-execution agent
4 of 5up to 5 re-execution tickets per day; it closed 4 out of 5 autonomously
#

What I learned / next step

  • >Memory persistence is a design problem, not a library choice: I prototyped with ChromaDB locally and migrated everything to DynamoDB for production. What an agent should remember, for how long and where to store it defines its behavior.
  • >Instructions and context are an agent's real code: I learned to refine the prompts and the context each specialist receives.
  • >Without evals, monitoring and observability there is no reliable improvement: every change must be measured and production behavior must be watched.
  • >An agent shouldn't be a copilot: the value is in self-managing the full loop (query, diagnose, re-execute) with a human supervising, not executing steps on request.