Multi-agent system for data operations and support
Agents that query, diagnose and re-run integration jobs in production.
Problem
Every customer came with different data sources, and every integration created recurring operational work: questions about job status, pipeline failures to diagnose by reading logs, and re-execution tickets that someone had to validate, launch and close by hand.
As an Integrations Engineer I was customer-facing. I listened to the problem in the meeting, understood the customer's source system and built the connection. After delivery, the cost was in daily operations: the same kind of ticket, several times a day, with the same diagnostic pattern.
The design question was: which part of that loop (query, diagnose, re-execute) can agents automate without losing control over customer data.
Architecture
A Router + Specialists pattern on a stateful graph (LangGraph, Google ADK):
flowchart TD
User["User Events (Slack / Zendesk / CLI)"] --> Router
subgraph EKS["Agent Services on AWS EKS"]
Router["AgentRouter (gemini-2.0-flash)"]
Router -->|"DATA_QUERY"| SQL["JobsSQLAgent"]
Router -->|"DOCUMENTATION"| Docs["DocsRAGAgent"]
Router -->|"DIAGNOSTIC"| Diag["SupportDiagAgent"]
Router -->|"REEXECUTION"| Reexec["ReexecutionAgent"]
end
SQL --> DB[("PostgreSQL via MCP")]
Docs --> KB["Markdown Knowledge Base"]
Diag --> DB
Diag --> CW[("AWS CloudWatch Logs")]
EKS -.-> DDB[("DynamoDB (State & Memory)")]
Reexec --> API[("Backend API")]
API --> Batch[("AWS Batch")]
classDef default fill:#0f172a,stroke:#38bdf8,stroke-width:1px,color:#e2e8f0
classDef router fill:#881337,stroke:#f43f5e,stroke-width:1px,color:#fecdd3
classDef specialist fill:#1e1b4b,stroke:#818cf8,stroke-width:1px,color:#e0e7ff
classDef storage fill:#022c22,stroke:#10b981,stroke-width:1px,color:#d1fae5
class Router router
class SQL,Docs,Diag,Reexec specialist
class DB,CW,API,Batch,KB,DDB storageKey decision: the re-execution agent has no query or diagnostic logic of its own. It calls the specialists, so each piece can be evaluated separately.
Results
Measured with my own evaluation, with no customer data exposed:
What I learned / next step
- >Memory persistence is a design problem, not a library choice: I prototyped with ChromaDB locally and migrated everything to DynamoDB for production. What an agent should remember, for how long and where to store it defines its behavior.
- >Instructions and context are an agent's real code: I learned to refine the prompts and the context each specialist receives.
- >Without evals, monitoring and observability there is no reliable improvement: every change must be measured and production behavior must be watched.
- >An agent shouldn't be a copilot: the value is in self-managing the full loop (query, diagnose, re-execute) with a human supervising, not executing steps on request.