Mastery · Track: data engineering · lesson 10.1 open
Where AI agents fit into data engineering
Real use cases, what infrastructure each requires, and where NOT to use an agent.
Promise: a map of data agent use cases—with risk, reward, and prerequisites—so you can choose your first one with clear criteria.
Why this is a differentiator
Almost every data engineer already uses coding assistants. Far fewer ship production agents on enterprise data, with evaluation, access control, and measured cost. That is the kind of experience verified in a 45-minute interview, because it requires architectural decisions and not just tools. Lesson 9.4 covers how to prove this; this module covers how to build it.
This module is conceptual and architectural. Tools and versions change fast; what lasts are the principles (contract, validation, evaluation, governance).
What an agent is, for this module
A program in which a language model decides which tools to call, and in what order, until a task is completed. In data, tools are typically queries, analytical APIs, catalog searches, and quality checks.
Use cases
| Use case | Input | Output | Risk | Prerequisite |
|---|---|---|---|---|
| Natural language questions about metrics (NL→SQL or NL→API) | analyst question | number or table + explanation | wrong answer that looks right | curated metric definitions; evaluation |
| Pipeline diagnostics | delay or failure alert | root-cause hypothesis + log links | incorrect hypothesis | read-only access to logs and metadata |
| Documentation and lineage | tables and jobs | descriptions and dependency map | incorrect description | human review before publishing |
| Quality triage | data test results | severity classification + suggested action | false negative | labeled historical data |
| Test and migration generation | schema and rules | code for review | plausible and incorrect code | automated validation + review |
Where NOT to use an agent
- An irreversible action without human confirmation (deleting, overwriting, sending).
- A financial or regulatory calculation that must be deterministic and auditable; use traditional code.
- When a standard API solves the problem (a filter, a fixed aggregation).
- When you cannot measure whether it was accurate.
A selection rule
Score from 0 to 2: (a) does the task have a verifiable answer? (b) is an error cheap? (c) is evaluation data available? (d) is the time saving recurring? Total ≥ 6: good candidate. Below that: automate with traditional code.
List five repetitive tasks from your data team, score them using the four questions, and choose the best candidate. Write in one sentence how you would know if the agent was accurate.
Checklist
- Did I choose the use case based on criteria, not novelty?
- Do I know how to measure accuracy?
Lessons cited by number that are not in the library yet open later.
Checked on 05/10/2026. Educational content; tax and legal: confirm with a professional.