Many data pipelines were never designed for the scale and reliability modern AI workloads demand.
To keep pace with change, organizations are moving toward autonomous data pipelines that can automatically detect issues, adapt to changing conditions, and ensure reliable delivery of trusted data for AI and analytics.
DBTA recently held a roundtable webinar, Building Autonomous Data Pipelines for the AI Era, with experts who explored how this shift is reshaping modern data architecture.
The gap between AI ambition and execution is the market opportunity and it’s wide open, said Vinay Bachappanavar, product manager, data integration at Informatica from Salesforce.
The key challenges data leaders need to solve to make enterprises agentic include:
Context: How do we ground AI agents with trusted data and unified business context?
Access: How do we deliver deep data management capabilities directly inside agent deployment environments?
Governance: Hoe do we ensure that the enterprise data policies are enforced in an agentic enterprise?
Bachappanavar introduced Informatica Headless Data Management—a capability that surfaces Informatica’s deep data management intelligence natively within any application, platform, or AI framework.
Paul Watson-Gover, senior solutions architect at WhereScape, defined data automation. According to The Data Warehouse Institute (TDWI), data platform automation is “…using technology to gain efficiencies and improve effectiveness in data warehousing processes. Data platform automation is much more than simply automating the development process. It encompasses all of the core processes of data warehousing including design, development, testing, deployment, operations, impact analysis, and change management.”
Today’s risks and challenges include increasing complexity and chaos; manual inefficiency and cost; and increasing complexity and chaos where it takes months or years to build—resulting in failure.
WhereSpace Data Automation accelerates the delivery and management of data infrastructure, Watson-Gover said.
It addresses key data challenges, including:
- Eliminates time-consuming manual coding and reduces errors.
- Accelerates project delivery from months/years to days/weeks.
- Lowers maintenance costs and simplifies ongoing management.
- Facilitates smooth, rapid platform migrations (e.g., to the cloud) with minimal risk.
Yesterday’s data architectures slow AI adoption in the enterprise, explained Dave Yaffe, CEO and co-founder at Estuary.
Autonomous pipelines should:
Create: Humans or Agents: say what you need and get a working flow. Build both pipelines and connectors from a prompt. Guardrails to ensure everything is validated and tested before it touches production.
Heal: Pipelines that heal before humans know there’s a problem. Failures root-caused, fixes drafted and sandbox-tested, a PR opened for human approval.
Reason: Intelligence inside pipelines. Transformations that use LLMs to classify, enrich, and route events as they flow, continuously.
Agents need a platform with guardrails, Yaffe noted. Estuary v2.0 offers the following:
- Parallelism to work at any scale
- Exactly-once and schema-enforced, end to end
- A runtime that scales, recovers, and evolves on its own
- Real-time to batch, one dial
- APIs that agents know how to drive and humans can review
“Agents don’t add reliability. They inherit it. Get the foundation right. Then let the pipeline create, reason, and heal,” Yaffe concluded.
For the full webinar, featuring a more in-depth discussion, Q&A, and more, you can view an archived version of the webinar here.