The company is not named here. The infrastructure described is real, including the AWS services and the scale involved. The same pattern applies to any product built on a comparable cloud provider, an observability platform, and a knowledge management tool such as Notion or Confluence.
AI agents are often demonstrated through coding tasks, research summaries, or chat interfaces. Production infrastructure work is a more demanding environment.
The information needed to answer one operational question may be distributed across cloud metrics, application telemetry, source code, ticketing systems, pull requests, databases, and internal documentation.
The difficult part is often not producing the analysis. It is collecting the right evidence, correlating it across systems, preserving the context, and producing something another engineer can verify.
I use AI agents to reduce this coordination overhead.
The agent does not operate production systems independently. It connects to approved tools, gathers evidence, follows a defined investigation process, and produces a report. I remain responsible for validating the evidence and deciding what action to take.
This article describes three workflows I use in real infrastructure and production engineering work:
- Building a consolidated production health report for a customer tenant.
- Turning weekly engineering work into a searchable operating record.
- Investigating slow API requests and the database queries behind them.
Each workflow is narrow, repeatable, connected to real systems, and designed around an existing engineering process.
1. Building a production health report for a tenant
A customer-facing SaaS platform may store operational information across several systems.
Database utilisation may be visible in Amazon Aurora. Application compute utilisation may be visible in Amazon ECS. Errors and latency outliers may live in an observability platform such as Datadog, Dynatrace, or Grafana Cloud. Historical investigation reports may live in a knowledge system such as Notion or Confluence.
In my case, the workflow connects Aurora, ECS, the observability system, and Notion.
An engineer investigating a tenant would otherwise have to move between these systems and manually reconcile different timelines.
The first workflow automates much of that collection process.
Start with a specific operational question
The workflow begins with a bounded question:
The investigation may have been triggered by a performance review, capacity concern, incident follow-up, or customer escalation.
I give the agent the tenant, the relevant database cluster and application service, the reporting period, and the expected output.
Collect the database evidence
The agent retrieves Aurora metrics at 5-minute granularity for the previous seven days.
The collected signals include measures such as:
- CPU utilisation
- Read and write IOPS
- Storage consumption
- Database connections
- Other cluster-specific capacity indicators
For each signal, the agent records minimum, maximum, and average utilisation.
The distinction is useful.
An average may show that a database was healthy for most of the week while hiding a short period of saturation. A maximum without context can make a brief spike look more significant than it really was.
The agent therefore retains both aggregate values and the timing of unusual behaviour.
Correlate the ECS service
The next step is to retrieve the corresponding ECS metrics for the API tasks serving the tenant.
Depending on the service, this can include:
- CPU utilisation
- Memory utilisation
- Task count
- Task restarts
- Service utilisation
The important step is correlation.
The workflow can ask whether:
- Database CPU increased at the same time as API latency.
- Application memory pressure appeared before task restarts.
- Database behaviour remained stable while application latency increased.
- A resource spike was momentary or sustained.
This is more useful than producing another collection of independent dashboards.
The objective is to construct a timeline across the stack.
Add errors and latency evidence
The agent then connects to the observability platform and retrieves telemetry for the same reporting window.
It looks for signals such as:
- Application errors
- Repeated warnings
- Timeouts
- Latency outliers
- Unusual request behaviour
- Events corresponding to infrastructure metric spikes
I treat the output as an evidence package rather than an automatic root-cause determination.
A latency increase may be associated with database pressure, an expensive request pattern, application-level contention, an external service, or a combination of these.
The agent organises the signals. I decide what they mean.
Publish a consistent report
After collecting database, compute, and observability evidence, the agent produces a structured report and publishes it to Notion through the Notion MCP connection.
A typical report contains:
- Scope and reporting period
- Aurora utilisation summary
- ECS utilisation summary
- Error and latency findings
- Important timestamps
- Correlated events
- Potential capacity or reliability concerns
- Questions requiring further investigation
- References to the underlying evidence
There are two benefits.
The obvious one is time. I do not have to manually collect the same information from several systems.
The second is consistency.
Tenant reviews performed several weeks apart use the same metric set, the same report structure, and the same investigation sequence.
2. Creating a weekly engineering operating record
The second workflow addresses a different problem.
Engineering work is fragmented across tickets, pull requests, reviews, debugging sessions, operational interventions, and discussions.
A ticket title rarely explains the full value of the work.
The associated pull request may contain the implementation details. Review comments may contain the architectural reasoning. Operational impact may only become obvious after looking at all of them together.
This information is useful during weekly reviews, performance discussions, knowledge transfer, planning, and retrospectives.
It is also expensive to reconstruct manually.
Linear is the index of work
Every meaningful unit of engineering work I do is associated with a Linear ticket.
A similar workflow could use another work-tracking system such as Jira, but Linear is the system I use.
I classify tickets using tags such as:
- Keep-the-lights-on work
- Technical debt
- Customer requests
- Debugging and troubleshooting
- Reliability improvements
- Platform engineering
- Operational maintenance
This discipline is important.
An AI agent cannot reliably reconstruct engineering work when that work is spread across undocumented conversations, local notes, and unrelated pull requests.
The ticket becomes the starting point from which the rest of the evidence can be discovered.
Retrieve the completed work
At the end of each week, I ask a Codex agent to retrieve the tickets closed during the previous week.
For each ticket it collects information such as:
- Ticket title and description
- Category and tags
- Completion date
- Linked pull requests
- Relevant comments
It then follows the implementation trail into GitHub.
Read what actually changed
Using the GitHub MCP connection, the agent inspects each pull request associated with the ticket.
It reviews:
- Pull request description
- Code changes
- Files modified
- Review comments
- Requested changes
- Follow-up discussions
- Final merge state
The purpose is to answer three questions:
- What problem was being solved?
- What changed in the system?
- What was the expected benefit?
This matters because a ticket title is often an incomplete description of the work.
A ticket may say "Improve deployment reliability."
The implementation may reveal that I changed retry behaviour, improved idempotency, added validation, or fixed a race condition.
Review comments may explain why the first implementation had a problem and how the final design addressed the risk.
Build the historical record in Notion
The weekly output is written into a structured Notion database.
Each record can contain:
- Completion date
- Linear ticket
- Work category
- Problem statement
- Implementation summary
- Pull request links
- Operational or customer impact
- Follow-up work
- Supporting evidence
Several useful views become available once the data is structured.
A calendar view shows what was completed during the month.
A category view shows how engineering effort was distributed across maintenance, customer requests, technical debt, reliability, and platform improvements.
Over time, this becomes a searchable operating record rather than a collection of disconnected tickets.
Which of the three workflows would you recommend building first?
Better performance context without reducing engineering to ticket counts
The record also provides useful context during performance discussions.
It should not become an automated employee scoring mechanism.
Ticket count is not a complete measure of engineering contribution. One difficult production investigation may require more skill and judgment than several small changes.
Architecture work, mentoring, prevention work, reviews, and cross-team support can also be poorly represented by simple activity metrics.
The useful information is the shape of the work.
A manager can examine questions such as:
- How much work addressed customer problems?
- How much reduced operational risk?
- Which issues required deep debugging?
- How much technical debt was addressed?
- Did the engineer repeatedly handle urgent operational work?
- Which changes improved developer or platform productivity?
The AI agent reconstructs and organises the evidence.
The manager and engineer interpret it.
3. Investigating slow API requests and database latency
The third workflow supports a technically demanding investigation.
A customer sends an incoming API request.
The request reaches the load balancer, which forwards it to the API service. The API processes the request, generates the database operations required to fulfil it, sends SQL to PostgreSQL, and returns a response to the customer.
Depending on what the customer requested, the resulting SQL may become complex.
Complexity can be affected by:
- The amount of data requested
- The number of models or entities involved
- Nested relationships
- Filtering and sorting
- Access-control conditions
- Tenant-specific data distribution
- The number of joins required to construct the response
When an API request is slow, I need to understand both the application request and the database behaviour it triggered.
Start with the observability evidence
During execution, the API records diagnostic information into the observability platform.
The telemetry can include information such as:
- The incoming API request
- Request components stored in compressed form
- Database execution duration
- Overall request latency
- Request identifiers
- Tenant identifiers
- Other execution context
Some of the request information is Brotli-compressed to reduce logging volume.
That is useful operationally, but it makes a manual investigation more cumbersome.
I would otherwise have to:
- Locate the correct log event.
- Extract the compressed information.
- Decompress it.
- Reconstruct the relevant request context.
- Understand how the API converted that request into database operations.
- Identify the SQL involved.
- Examine how PostgreSQL planned it.
This is a good candidate for an agent because several mechanical steps have to be completed before the engineering analysis can begin.
Use the application codebase as evidence
A Codex task performs much of the reconstruction.
It retrieves the relevant telemetry from the observability platform.
It extracts and decompresses the required request information.
It then uses the API codebase to trace how that request is transformed into the SQL sent to PostgreSQL.
Access to the codebase is important.
Once the SQL has been reconstructed, the workflow examines it using the PostgreSQL planner.
Analyse the database plan
The agent can identify signals such as:
- Sequential scans over large relations
- Missing or ineffective indexes
- Poor join order
- Expensive joins
- Large intermediate result sets
- Cardinality estimation errors
- Expensive sorting or aggregation
- Tenant-specific data skew
- Excessive data requested by the API call
- A large difference between database time and total request latency
The final report connects the request-level evidence to the database-level evidence.
For example, the report may show that a particular API request resulted in several joins over large relations, that PostgreSQL substantially underestimated the number of rows involved, and that most database execution time was concentrated in one part of the plan.
That gives me a much stronger starting point than simply knowing that the API was slow.
A strong diagnostic signal, not an autonomous conclusion
This workflow is not completely self-contained.
A database execution plan does not always explain total customer-facing latency.
Other contributors can include:
- Connection pool contention
- Locks
- Application processing
- Network latency
- Serialization
- Cache behaviour
- External services
- Resource contention elsewhere in the stack
The agent therefore produces a strong diagnostic signal rather than a final verdict.
I review the evidence and decide whether the next step should be an index change, SQL change, application change, API restriction, statistics refresh, capacity adjustment, or additional measurement.
The agent performs the mechanical reconstruction and first-pass analysis.
I own the operational conclusion.
What these three workflows have in common
The problems are different, but the agent design is similar.
1. The task has a boundary
Each workflow has a known starting point and a known output.
I do not ask the agent to "manage production" or "improve performance."
I ask it to:
- Analyse one tenant for a specific time window.
- Summarise a defined set of completed tickets.
- Investigate one slow API request.
Boundaries make agent behaviour easier to test, reason about, and audit.
2. The agent works from source systems
The workflows retrieve evidence from real systems: Aurora, ECS, an observability platform, Linear, GitHub, PostgreSQL, and Notion.
The agent does not need to answer from general knowledge when production evidence exists.
This also makes the result more verifiable because the report can retain references to the original metrics, telemetry, tickets, pull requests, and database plans.
3. The workflow follows an existing engineering process
The agent is effective because the investigation process is already understood.
I already know which metrics matter in a tenant review.
I already understand how Linear tickets relate to GitHub pull requests.
The application already defines how incoming requests become database operations.
The agent executes these known processes more consistently.
It does not replace the need to define them.
4. The result becomes organisational memory
Useful agent output should not disappear inside a chat session.
Publishing reports into Notion makes them searchable, comparable, and accessible to the team.
A one-time AI response becomes part of the engineering record.
5. A human remains accountable
None of these workflows requires an agent to have unrestricted authority to modify production.
The agent gathers evidence, transforms information, identifies signals, and prepares an analysis.
Controls for production AI agents
Similar workflows should begin with narrow permissions and explicit boundaries.
Useful controls include:
- Read-only access wherever possible
- Least-privilege credentials
- Tenant and environment scoping
- Approved query and command patterns
- Limits on metric and telemetry retrieval
- Sensitive-data redaction
- Audit trails for tool calls
- References back to original evidence
- Human approval before production changes
- Clear failure and escalation behaviour
What access does the agent actually need?
Database investigations require additional care.
An agent analysing query performance should not automatically receive permission to execute arbitrary mutations. Diagnostic queries and planner operations should use read-only access, safe environments, replicas, or other controls appropriate to the production architecture.
Engineering performance information also requires judgment. Agent-generated activity summaries can provide context, but they should not be treated as objective employee scores.
Where these use cases fit in an AI strategy
In the Tech Continuum AI Pattern Selector, these workflows fall mainly under Data Processing and Decisioning.
The agents collect information from several systems, transform it into a consistent representation, identify relevant signals, and prepare evidence for an engineering decision.
That describes the type of AI workload.
Production readiness is a separate question.
An organisation also needs suitable access controls, data quality, governance, evaluation, operational ownership, security boundaries, and processes for responding when the agent is wrong.
The Tech Continuum 7-Pillar AI Readiness Scorecard examines that broader operating environment.
The two frameworks answer different questions.
The AI Pattern Selector helps classify what you are building.
How does this relate to Tech Continuum's AI readiness framework?
The practical value of production AI agents
The main benefit in these workflows is not autonomous infrastructure management.
It is the removal of repeated coordination work.
The first agent gathers and correlates information that would otherwise require several systems and several manual steps.
The second reconstructs engineering work from tickets, pull requests, code, and review discussions.
The third retrieves production telemetry, decompresses request context, follows application code, reconstructs SQL, and prepares a first-pass database analysis.
I spend less time assembling evidence and more time evaluating it.
The process also becomes more consistent. The same signals are collected. The same investigation sequence is followed. The same source evidence is retained.
That is where I currently find AI agents most useful in production infrastructure engineering.
Do I need a dedicated AI platform team to build workflows like these?
Frequently asked questions
No. All three were built by one infrastructure engineer, using tools already in place. The barrier is not platform investment. It is knowing your own workflow well enough to hand a bounded piece of it to an agent.
No. These are internal tools, not customer-facing features. The AI Pattern Selector classifies customer-facing feature work. This is a different category: using an agent as a force multiplier on infrastructure work I already do.
Read-only, in all three cases. Metrics access to Aurora and ECS, telemetry access to the observability platform, ticket access to Linear, and repository access to GitHub. None of these tasks write to production systems.
The weekly engineering operating record. It has the clearest boundary, the lowest risk if something goes wrong, and it produces something genuinely useful within a week.
The 7-Pillar framework's Integration, MLOps, and Lifecycle pillar is the closest match. The same discipline that makes a customer-facing AI feature safe in production, bounded scope, monitored output, a human in the loop, applies here too, just pointed at internal tooling instead of a product feature.