In this interview from Snowflake Summit 2026 in San Francisco, Vaibhav Jajoo, head of data engineering, data platform and business intelligence at DoorDash, joins Chris Child, vice president of product management at Snowflake, to talk with theCUBE's Dave Vellante and Rebecca Knight about how a unified, governed data foundation enables the shift from experimental AI to production-ready agentic workflows. Jajoo describes how DoorDash processes petabytes of data daily across consumers, Dashers and merchants, and explains why machine users — including ML features and agentic workflows — now outpace human analysts in data consumption. He details how adopting an open architecture with Iceberg across its full data estate has cut costly data movement, reduced latency and freed engineers to focus on business outcomes rather than pipeline plumbing. Child adds that companies deploying agents most effectively are those that have built well-defined, well-documented data products rather than simply exposing thousands of tables.
The conversation also explores how data engineering roles are evolving — not collapsing — as AI handles repetitive tasks and teams increasingly cross traditional boundaries. Child introduces Snowflake's CoCo and CoWork tools, designed respectively for teams building governed data products and teams consuming them, reflecting the platform's push to support all three layers of the modern data organization. Jajoo unpacks how DoorDash is constructing learning loops that teach agents which assets to invoke in a given context, preventing the confident hallucinations that arise when agents navigate an undifferentiated data lake. Child underscores Snowflake's strategic decision to bring models to the data rather than route governed, row-level access controls through external LLM integrations. From DoorDash's customer-first competitive philosophy to early-stage robotics experiments with its delivery robots, the discussion makes clear that durable AI outcomes begin with disciplined data fundamentals.
Find more SiliconANGLE news and analysis https://siliconangle.com/
Follow theCUBE's wall-to-wall event coverage https://siliconangle.com/events/
Learn about the latest theCUBE events https://www.thecube.net/
00:00 - Intro
00:06 - Introduction to DoorDash: Welcome and Data Architecture Overview
02:13 - Snowflake: Strategies in Data Engineering and Governance
04:35 - Advancing Data Engineering: Evolution and Roles at DoorDash
09:13 - Unified Data Strategy: Collaboration and Consistency
11:44 - AI-Driven Workflows: Transforming Business and Data Utilization
14:27 - Competitive Strategy and Innovation at DoorDash
16:42 - Innovative Collaboration and AI Integration: The Future at DoorDash
#theCUBE #SnowflakeSummit #theCUBEresearch #Snowflake #DoorDash #AI #DataEngineering
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Snowflake Summit 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Register for Snowflake Summit 2026
Please fill out the information below. You will receive an email with a verification link confirming your registration. Click the link to automatically sign into the site.
You’re almost there!
We just sent you a verification email. Please click the verification button in the email. Once your email address is verified, you will have full access to all event content for Snowflake Summit 2026.
I want my badge and interests to be visible to all attendees.
Checking this box will display your presense on the attendees list, view your profile and allow other attendees to contact you via 1-1 chat. Read the Privacy Policy. At any time, you can choose to disable this preference.
Select your Interests!
add
Upload your photo
Uploading..
OR
Connect via Twitter
Connect via Linkedin
EDIT PASSWORD
Share
Forgot Password
Almost there!
We just sent you a verification email. Please verify your account to gain access to
Snowflake Summit 2026. If you don’t think you received an email check your
spam folder.
In order to sign in, enter the email address you used to registered for the event. Once completed, you will receive an email with a verification link. Open the link to automatically sign into the site.
Sign in to gain access to Snowflake Summit 2026
Please sign in with LinkedIn to continue to Snowflake Summit 2026. Signing in with LinkedIn ensures a professional environment.
Are you sure you want to remove access rights for this user?
Details
Manage Access
email address
Community Invitation
Chris Child, Snowflake & Vaibhav (VJ) Jajoo, DoorDash
In this interview from Snowflake Summit 2026 in San Francisco, Vaibhav Jajoo, head of data engineering, data platform and business intelligence at DoorDash, joins Chris Child, vice president of product management at Snowflake, to talk with theCUBE's Dave Vellante and Rebecca Knight about how a unified, governed data foundation enables the shift from experimental AI to production-ready agentic workflows. Jajoo describes how DoorDash processes petabytes of data daily across consumers, Dashers and merchants, and explains why machine users — including ML features and agentic workflows — now outpace human analysts in data consumption. He details how adopting an open architecture with Iceberg across its full data estate has cut costly data movement, reduced latency and freed engineers to focus on business outcomes rather than pipeline plumbing. Child adds that companies deploying agents most effectively are those that have built well-defined, well-documented data products rather than simply exposing thousands of tables.
The conversation also explores how data engineering roles are evolving — not collapsing — as AI handles repetitive tasks and teams increasingly cross traditional boundaries. Child introduces Snowflake's CoCo and CoWork tools, designed respectively for teams building governed data products and teams consuming them, reflecting the platform's push to support all three layers of the modern data organization. Jajoo unpacks how DoorDash is constructing learning loops that teach agents which assets to invoke in a given context, preventing the confident hallucinations that arise when agents navigate an undifferentiated data lake. Child underscores Snowflake's strategic decision to bring models to the data rather than route governed, row-level access controls through external LLM integrations. From DoorDash's customer-first competitive philosophy to early-stage robotics experiments with its delivery robots, the discussion makes clear that durable AI outcomes begin with disciplined data fundamentals.
Find more SiliconANGLE news and analysis https://siliconangle.com/
Follow theCUBE's wall-to-wall event coverage https://siliconangle.com/events/
Learn about the latest theCUBE events https://www.thecube.net/
00:00 - Intro
00:06 - Introduction to DoorDash: Welcome and Data Architecture Overview
02:13 - Snowflake: Strategies in Data Engineering and Governance
04:35 - Advancing Data Engineering: Evolution and Roles at DoorDash
09:13 - Unified Data Strategy: Collaboration and Consistency
11:44 - AI-Driven Workflows: Transforming Business and Data Utilization
14:27 - Competitive Strategy and Innovation at DoorDash
16:42 - Innovative Collaboration and AI Integration: The Future at DoorDash
#theCUBE #SnowflakeSummit #theCUBEresearch #Snowflake #DoorDash #AI #DataEngineering
Chris Child, Snowflake & Vaibhav (VJ) Jajoo, DoorDash
Chris Child
VP of Product, Data EngineeringSnowflake
Vaibhav (VJ) Jajoo
Head of Data Engineering, Data Platform & BIDoorDash
search
Rebecca Knight
>> Good morning, everyone, and welcome back to theCUBE's live coverage of the Snowflake Summit 2026 here at the Moscone Center in San Francisco. We are in day two. I'm your host, Rebecca Knight, alongside Dave Vellante, the co-founder of theCUBE. I would like to welcome two fantastic guests to the show, Vaibhav Jajoo, head of data engineering and data platform at BI at DoorDash. Welcome, VJ.>> Thank you so much.
Rebecca Knight
>> And Chris Child, VP of product and data engineering at Snowflake. Welcome.>> Welcome. Thank you.
Rebecca Knight
>> Welcome back to the show, I should say.>> Thank you. Thank you.
Rebecca Knight
>> VJ, DoorDash single-handedly feeds my children, but it's also one of the largest real-time logistics companies in the world. Why don't you walk us through how you have built the data foundation to handle both the real-time side and also the analytical side?>> Thank you so much for the business, first of all. I really appreciate it.
Dave Vellante
>> Best service out there. I would say that right now. If you haven't tried DoorDash, it's amazing.>> Thank you so much for your business. Really appreciate it. So what we have learned over time is, when you are operating DoorDash, which is one of the world's largest logistics platform processing petabytes of data on a daily basis, we have to make sure that we do it right for both consumer, Dashers, and the merchants. So we are collecting information from all of them, and you have to do that with a high quality, low latency while making sure that you're doing it efficiently. Over the last 10 years, we've learned that it's super important we adopt a platform which is open architecture, open storage, and open compute. That then creates a very interesting pay for what is changing in terms of use cases and workload. What we have learned over time is that the machine user is outpacing the human user in consumption of analytics data. And that is very interesting because the ML features, the feedback loops to the production services or the AI agentic workflows are outpacing the analytics user. So when you do that, you cannot adopt a monolithic environment which is holding you back and not letting you enable new use cases on top of it. So open architecture, which is compute agnostic, which is storage agnostics, helps us create an architecture which is easy to scale and consume across.
Dave Vellante
>> Chris, in the analyst session the other day, you were talking about how Iceberg is obviously central to your strategy now, how you're trying to minimize data movement, how it's changing data engineering. We were talking to VJ about how DoorDash, and we'll get into this, is serving not just the consumer, but the merchants and the Dashers. So all of that sort of complexity falls on your data shoulders, so tie it into some of the comments you're making->> One of the things is, I think it's been fantastic to be close partners with VJ and his team for many years at this point. One of the problems they came to us with, I would say, years ago at this point was that, like many companies, they don't only use Snowflake. They have other data platforms, they have other tools that they want to be able to use. Some of those are for... They have a team that really focuses on how they optimize the Dasher experience, a team that really focuses on how they optimize the consumer experience. In some cases, they were using different tools to do that job. And they were spending a bunch of money moving data between those different tools, bringing data into Snowflake, bringing data out of Snowflake. Each of those hops is expensive, and frankly, kind of unnecessary. So as we started investing heavily in Iceberg, we were able to help them adopt Iceberg across not just Snowflake, but their much broader data estate as well. The end result has been that it's cheaper, it's faster, they have lower latency on data because there's not these movement steps in between. And the data engineers are able to focus, as we've talked about before, on solving business problems rather than building plumbing and moving data around. It's really the Iceberg investment that we've made together that's enabled that.
Dave Vellante
>> VJ, how do you think about the governance piece of that? Because a couple years ago, when Iceberg was starting to gain momentum, you would ask practitioners like yourself, "Well, how are you going to govern it?" And they didn't have a good answer. So two years on, what's your answer?>> Well, my perspective on the governance is very simple on this. You create data in one standard way, and you open up for multiple use cases. You govern at the storage layer, then you're able to apply that same governance across all of the platforms. You define the data quality principles in one central way. You come back and think about compute optimization, authorization in one central way, and then you are able to manage the SLOs as well as the quality of the data at the same time. Makes it super easy for us, standardized structure.
Dave Vellante
>> So you get that consistency.
Dave Vellante
>> It makes it super easy for us. Exactly.
Rebecca Knight
>> So talk a little bit about how data engineering has evolved within DoorDash, and maybe just compare it today to what it was, say, a few years ago.>> What is interesting actually, if you look at it from an agentic workflow standpoint, that has been the latest advent of it. What we have learned is, when you are producing data, you cannot forget about the basic data engineering principles, which is thinking about high-quality data assets, which are reliable. They are meeting the latency and the SLO guarantees that they're looking for. And then making sure that they're consistently available to every workflow that you use. If you think about as an example for the agentic workflow, you cannot throw out thousands of tables out there for people to consume from or an agent to consume from without a good proper semantic around it, good context model around it, or good documentation around it. So the fundamentals of data engineering has not changed, but it is adopted to the new reality, and coming back and saying that, "Hey, how do I produce a high-quality asset which is documented well, which can be used by the ML feature, which could be used by your product loop back systems, or which could be used by agentic platform seamlessly?" So that makes it simple, easy.
Dave Vellante
>> I've been asking this question for months and sometimes I get different answers. The Hadoop world created all these very highly specialized roles, data engineers, data analysts, quality analysts, business analysts, on and on and on. Do those roles, do you see them collapsing? Or are they just giving superpowers to those specialists? How do you think about that?>> I think of the latter more than anything else, because every job has a mundane activity or repetitive activity. What we are seeing is that repeated activity can be automated while you can go out then build on the second layer of the, what should I say, oomph on top of it. Data engineers are more productive with your AI-based development for it, but now they are focusing on producing better outcomes, improving the quality, making the workflows become better. So that's what is happening right now. Across the profiles, they're not collapsing in one particular thing, but they're progressively becoming better with each other. But interestingly, what is also happening at the same time is people are crossing the boundaries. They're not defining themselves as, "I'm at this persona versus I'm that persona." Because people can easily travel from one persona to another persona, because they can do the mundane of what I was waiting for somebody else to do for me till yesterday. Because I know the code base, I can go out there and make the change myself rather than waiting for somebody to create a ticket and build the system online.
Dave Vellante
>> So you can scale that capability without necessarily scaling the team. Is it fair to say that those mundane activities are largely coordination kind of paper cuts and the judgment piece of that, you called it, I think, strategic, is left to the humans? But square this circle for me, VJ, because you also said something that makes a lot of sense. I can now traverse those different roles because I have an infinitely intelligent and patient AI partner that can help me do that. So why aren't those things in conflict, what you just said about traversing and the collapsing of those roles?>> See, that's how I think about it. There are some specialist actions, let's say, talk about very deep analysis on a real-time system, coming back and thinking about the latency for it. It's a data engineer and data platform engineer's job. They will do it most effectively. They'll understand how to think about it in an efficient way, a reliable way. That will be the deep stack work that a data engineer will most probably do while exposing them to an analyst organization, let's say, for example, to come back and consume from that. That was something that the data engineer was spending a lot of time with. Now, if that second job is more self-serve for rest of the organization, then they can come back and investigate what's happening with it. So the true optimization is what the data engineer is focusing on.
Dave Vellante
>> Chris, these are your peeps, and so you've seen the evolving role of the data engineer. Do you have any thoughts on this topic?
Rebecca Knight
>> Because it's a real cultural shift within an organization.>> Yeah, sure. Well, I don't know, I'll tie this to announcements, but this is a lot of why we've focused a lot of our energy going forward on CoCo and CoWork. I think it actually fits very well into the way you were just describing this. CoCo is a set of tools to help the people who are building these data products, building the data pipelines, really getting the data ready to be consumed. And then CoWork is the platform for the people who are consuming the data. So I really think of it as there's three roles as this sort of evolves. I actually do think the names and definitions of these roles will evolve. They won't disappear and completely change, but they will evolve and look different, and we may call them different things down the line. But there are other people who are making sure that the data is in the right shape, is usable, is trustable, the right semantic models are in place. The data is set up well to be consumed by AI or by a person. And then you have the people who are going to take that data and are building the tools and agents and capabilities that your business users, the third group that you care about, are going to be able to come in and consume. But I think it will shift because the business users are no longer going to ask someone to write a query for them. Instead, they're going to ask an agent, and the data scientists will evolve to be getting the agents ready. And then you have the team that still needs to make sure the data is ready for them, but they're going to be orchestrating and managing teams of agents and other things to actually manage those pipelines. But they still need to be thinking, as you said, about what are the right pipelines? What do they cost? Do I have the right governance in place? You still need human judgment in all this.
Rebecca Knight
>> Absolutely.
Dave Vellante
>> If I go back to the Gartner variety, I'll go back there for a moment, you got structured data, unstructured data, you got streaming, you guys have announced some new Kafka capabilities, et cetera. So you've got to... I mean, the data engineer, the quality engineer, they spent a lot of time wrangling and reconciling all of that. I know you can, but specifically, how are you using R&D and product and technology to attack that problem with cross-lakehouse capabilities, et cetera?
Dave Vellante
>> So I actually think a lot of the problems that people run into coming from having different silos and different pipelines and different places that they're building and managing data. So you'll end up with multiple different teams who have each defined a certain metric in slightly different ways. That's where a lot of the wrangling and a lot of having to reconcile different data sources comes from. Generally speaking, if you go all the way back to a source system, in most cases there is ultimately a source of truth. There is a true fact that is happening somewhere in a system. And if not, if your data is entered incorrectly in Salesforce or wherever, you should fix it in the source system. Then what you need to do is have a single place that you're defining the downstream metrics and definitions that you need. So if you want to ask a question about what do we know about this Dasher, as an example, there should be a single definition of that, and that means that the place that owns that definition needs to have all of the data about that Dasher. It needs to have access to the unstructured data, the structured data, the streaming data, everything funneling into that one place so that you can have a single source of truth on what do you know about that particular fact.>> And then single source of truth is so important actually, because any use case that we are talking about in the infrastructure should talk to the same exact asset. You cannot hallucinate about it. You can go into multiple layers, go back down to the stack and get that data out, but that'll be confusing. Everybody will come back out with the wrong and incorrect information about it. Now, if you have one on think about that Dasher and have a proper information, then agentic workflow, ML workflows, or everybody should come back to the exact same system of record and make the same exact decision, so that the human and machines are taking the same exact decisions.
Dave Vellante
>> So when your agent encounters conflicting data, does the system say, "Okay, you got conflicting data, is it A or B?" and you reconcile it at that point, bring the human into the loop? First part of the question. And the second part of the question is, does the agent learn from the reasoning traces of the human and the judgment that it made so that the next time it encounters a similar conflict, it reconciles it without the human? Lond-winded question, but that's the nirvana.>> That's the hard part actually. If you expose thousands of tables to an agentic workflow and expect it to make the right call for it, it's going to make agentic workflow come back with a response and make it very confidently. Because on a keyword, it picked up a table, thought that that's the perfect information or perfect source for it, and came back with a hallucinated result and very confidently say that, "Okay, this is the number." But the point is actually about exposing the right assets, the right quality guardrails, light context and semantics around it through agentic workflows so that it understand what's the right data assets to use. Now, how do you do that actually? Which is the second part of your question.
Dave Vellante
>> Yes, right.>> How does a human analyst who is used to querying a data in a particular way for deeper analytics versus for executive reporting, so to speak, how does the agent learn that? So we are creating learning loops for the agents to understand what are the right assets, what are the context behind it, how to use them, and for in both situation, which table has to be invoked versus another. So the operator who's looking for a very detailed information gets to a lower level fact table, versus an executive who's looking for an highly aggregated data, goes to that aggregated table which has the very well-defined metrics for it. So if you don't teach what is what, then agent will come back with an incorrect response.
Dave Vellante
>> This is next-level stuff. I mean, this is relatively recent that you're actually able to do this. How has it impacted your business?
Rebecca Knight
>> People are loving it, actually. We are looking at the productivity gain. People are excited about it. Actually, what I'm seeing right now is, people are saying, "Okay, I can do all this stuff by myself." Our executives will come back and say, "I want to resolve the Dasher problem directly myself because I got this email on that." So that's super exciting to see them do it.
Rebecca Knight
>> Chris, when you're hearing VJ, I mean, obviously you work closely with VJ and have great rapport, but this is one of the largest logistics companies in the world. What would you say are some of the best practices that you are learning that you are also bringing to other customers about the success that DoorDash has had?>> It's a lot of what VJ's been talking about. I think what we're seeing repeatedly is, the companies who are able to build the best agents and deploy them are doing that because they've been very thoughtful about what data products those agents should be depending on, and those data products are well-defined. Companies that have thousands of tables in a gold layer and they try to put agents on top, it's actually very difficult because they run into points like what Dave was asking about, which is the right answer, it's not clear. Where you're much more prescriptive about these are the correct facts about Dashers, these are the correct facts about consumers, these are the correct facts about merchants, the agent is able to then reason much more effectively about it. So this is part of the... What we've taken from them is having the data engineering team elevate to be defining those facts and defining that. And facts, not really in the sense of like a fact or dim table, but really, what is the source of truth about these important pieces of information we have? Those are the ones who are able to build and deploy agents incredibly quickly. So we spend a lot of our time using the examples of customers like DoorDash who have been able to do this, and then trying to help the rest of our customer base get to that point because there is work and investment that you have to do to get to that place.
Dave Vellante
>> I want to ask you, VJ. We were in the analyst and journalist session with Sridhar. We asked them an unanswerable question. It was like, you've got the LLM vendors here, you got the SaaS vendors, they're all trying to own the system of intelligence. Obviously, Snowflake is heading in that direction. How do you see it playing out? And Sridhar said, "Look, I don't know the answer to that. Whoever wins is going to look back and rewrite history and say, 'We knew this was going to happen that way.'" It was very prescient. But what he did say is that, "Here's what I know: Product market fit is a life force, and so we need to innovate. That's how we're going to compete." So I thought that was a great answer. You were facing similar sort of competitive realities. People see what DoorDash does. Amazon is trying to get drones immediately, and you got Uber Eats out there, you got Walmart obviously competing, et cetera, et cetera, et cetera. How do you think about... Because you can't forecast the future, you don't know if drones and robots are going to be there, you obviously can take advantage of it well. So what's your mindset, your first principle on how you serve customers, Dashers, and your merchants?>> Our motto is one simple thing: Forget about the competition and think about making the product better for the three of them, for the consumer, for Dasher, and the merchant. So that's about innovation, thinking about where our product is off, and building it the right way. That has been our... I don't think we think about, as much as other companies do, about what is a competitor doing. Instead, we think about how can we make our product better, thinking about what capabilities we should build on. That's how we think about it.
Dave Vellante
>> No, that's the right answer. I mean, I think that's the mindset that Snowflake has based on what Sridhar said. I mean, because you can't predict what the competition is going to do. You can only sort of control what you can control.
Dave Vellante
>> It's innovation piece.>> Our philosophy is to pay very close attention to what our customers are doing, the problems they're running into, what they're trying to solve. Often a lot of, for us, the product development process is, look at what folks at the cutting edge like VJ are doing and then figure out how can we take that and turn it into a product that we can now sell to the rest of the customer base. That's been very successful for us of a way to find, what can we help them do? Let's work incredibly closely together and then let's go take that to everyone else.
Dave Vellante
>> But in the last three years, Chris, you guys have decided that we're not just going to be a data feed into these other intelligence systems, we're actually going to participate and add value to those intelligent systems. So that's a big strategic move that you guys->> I will say it came out of a similar thing where we saw these models come up, and we sat down with a number of customers, and they were saying, "Hey, we want to be able to use these not just with documents we have, but with all of the data that lives inside Snowflake. And we want to be able to do that in a high-trust, highly-governed way." And we tried to do that as an integration, and it didn't work very well. Part of that was because it didn't have the right context, it wasn't running within the context of the user. Companies do a lot of work around which specific rows can be accessed by which specific people for which specific purpose, and it was very hard to get that out to an LOM. So we found that the only way that we could really provide that same governance and high-quality answers to questions was to basically bring the models to the data rather than the other way around.
Dave Vellante
>> So many opportunities that will be created as a result.
Rebecca Knight
>> Indeed. It's exciting times. Well, VJ and Chris, thank you both so much for coming on the show. Really interesting conversation.
Dave Vellante
>> Thank you.>> Thank you so much. Really appreciate it.>> Thank you both.
Rebecca Knight
>> I'm Rebecca Knight for Dave Vellante. Stay tuned for more of day two of theCUBE's live coverage of the Snowflake Summit. You're watching theCUBE, the leader in enterprise tech news and analysis.
Chris Child, Snowflake & Vaibhav (VJ) Jajoo, DoorDash
search
Rebecca Knight
>> Good morning, everyone, and welcome back to theCUBE's live coverage of the Snowflake Summit 2026 here at the Moscone Center in San Francisco. We are in day two. I'm your host, Rebecca Knight, alongside Dave Vellante, the co-founder of theCUBE. I would like to welcome two fantastic guests to the show, Vaibhav Jajoo, head of data engineering and data platform at BI at DoorDash. Welcome, VJ.>> Thank you so much.
Rebecca Knight
>> And Chris Child, VP of product and data engineering at Snowflake. Welcome.>> Welcome. Thank you.
Rebecca Knight
>> Welcome back to the show, I should say.>> Thank you. Thank you.
Rebecca Knight
>> VJ, DoorDash single-handedly feeds my children, but it's also one of the largest real-time logistics companies in the world. Why don't you walk us through how you have built the data foundation to handle both the real-time side and also the analytical side?>> Thank you so much for the business, first of all. I really appreciate it.
Dave Vellante
>> Best service out there. I would say that right now. If you haven't tried DoorDash, it's amazing.>> Thank you so much for your business. Really appreciate it. So what we have learned over time is, when you are operating DoorDash, which is one of the world's largest logistics platform processing petabytes of data on a daily basis, we have to make sure that we do it right for both consumer, Dashers, and the merchants. So we are collecting information from all of them, and you have to do that with a high quality, low latency while making sure that you're doing it efficiently. Over the last 10 years, we've learned that it's super important we adopt a platform which is open architecture, open storage, and open compute. That then creates a very interesting pay for what is changing in terms of use cases and workload. What we have learned over time is that the machine user is outpacing the human user in consumption of analytics data. And that is very interesting because the ML features, the feedback loops to the production services or the AI agentic workflows are outpacing the analytics user. So when you do that, you cannot adopt a monolithic environment which is holding you back and not letting you enable new use cases on top of it. So open architecture, which is compute agnostic, which is storage agnostics, helps us create an architecture which is easy to scale and consume across.
Dave Vellante
>> Chris, in the analyst session the other day, you were talking about how Iceberg is obviously central to your strategy now, how you're trying to minimize data movement, how it's changing data engineering. We were talking to VJ about how DoorDash, and we'll get into this, is serving not just the consumer, but the merchants and the Dashers. So all of that sort of complexity falls on your data shoulders, so tie it into some of the comments you're making->> One of the things is, I think it's been fantastic to be close partners with VJ and his team for many years at this point. One of the problems they came to us with, I would say, years ago at this point was that, like many companies, they don't only use Snowflake. They have other data platforms, they have other tools that they want to be able to use. Some of those are for... They have a team that really focuses on how they optimize the Dasher experience, a team that really focuses on how they optimize the consumer experience. In some cases, they were using different tools to do that job. And they were spending a bunch of money moving data between those different tools, bringing data into Snowflake, bringing data out of Snowflake. Each of those hops is expensive, and frankly, kind of unnecessary. So as we started investing heavily in Iceberg, we were able to help them adopt Iceberg across not just Snowflake, but their much broader data estate as well. The end result has been that it's cheaper, it's faster, they have lower latency on data because there's not these movement steps in between. And the data engineers are able to focus, as we've talked about before, on solving business problems rather than building plumbing and moving data around. It's really the Iceberg investment that we've made together that's enabled that.
Dave Vellante
>> VJ, how do you think about the governance piece of that? Because a couple years ago, when Iceberg was starting to gain momentum, you would ask practitioners like yourself, "Well, how are you going to govern it?" And they didn't have a good answer. So two years on, what's your answer?>> Well, my perspective on the governance is very simple on this. You create data in one standard way, and you open up for multiple use cases. You govern at the storage layer, then you're able to apply that same governance across all of the platforms. You define the data quality principles in one central way. You come back and think about compute optimization, authorization in one central way, and then you are able to manage the SLOs as well as the quality of the data at the same time. Makes it super easy for us, standardized structure.
Dave Vellante
>> So you get that consistency.
Dave Vellante
>> It makes it super easy for us. Exactly.
Rebecca Knight
>> So talk a little bit about how data engineering has evolved within DoorDash, and maybe just compare it today to what it was, say, a few years ago.>> What is interesting actually, if you look at it from an agentic workflow standpoint, that has been the latest advent of it. What we have learned is, when you are producing data, you cannot forget about the basic data engineering principles, which is thinking about high-quality data assets, which are reliable. They are meeting the latency and the SLO guarantees that they're looking for. And then making sure that they're consistently available to every workflow that you use. If you think about as an example for the agentic workflow, you cannot throw out thousands of tables out there for people to consume from or an agent to consume from without a good proper semantic around it, good context model around it, or good documentation around it. So the fundamentals of data engineering has not changed, but it is adopted to the new reality, and coming back and saying that, "Hey, how do I produce a high-quality asset which is documented well, which can be used by the ML feature, which could be used by your product loop back systems, or which could be used by agentic platform seamlessly?" So that makes it simple, easy.
Dave Vellante
>> I've been asking this question for months and sometimes I get different answers. The Hadoop world created all these very highly specialized roles, data engineers, data analysts, quality analysts, business analysts, on and on and on. Do those roles, do you see them collapsing? Or are they just giving superpowers to those specialists? How do you think about that?>> I think of the latter more than anything else, because every job has a mundane activity or repetitive activity. What we are seeing is that repeated activity can be automated while you can go out then build on the second layer of the, what should I say, oomph on top of it. Data engineers are more productive with your AI-based development for it, but now they are focusing on producing better outcomes, improving the quality, making the workflows become better. So that's what is happening right now. Across the profiles, they're not collapsing in one particular thing, but they're progressively becoming better with each other. But interestingly, what is also happening at the same time is people are crossing the boundaries. They're not defining themselves as, "I'm at this persona versus I'm that persona." Because people can easily travel from one persona to another persona, because they can do the mundane of what I was waiting for somebody else to do for me till yesterday. Because I know the code base, I can go out there and make the change myself rather than waiting for somebody to create a ticket and build the system online.
Dave Vellante
>> So you can scale that capability without necessarily scaling the team. Is it fair to say that those mundane activities are largely coordination kind of paper cuts and the judgment piece of that, you called it, I think, strategic, is left to the humans? But square this circle for me, VJ, because you also said something that makes a lot of sense. I can now traverse those different roles because I have an infinitely intelligent and patient AI partner that can help me do that. So why aren't those things in conflict, what you just said about traversing and the collapsing of those roles?>> See, that's how I think about it. There are some specialist actions, let's say, talk about very deep analysis on a real-time system, coming back and thinking about the latency for it. It's a data engineer and data platform engineer's job. They will do it most effectively. They'll understand how to think about it in an efficient way, a reliable way. That will be the deep stack work that a data engineer will most probably do while exposing them to an analyst organization, let's say, for example, to come back and consume from that. That was something that the data engineer was spending a lot of time with. Now, if that second job is more self-serve for rest of the organization, then they can come back and investigate what's happening with it. So the true optimization is what the data engineer is focusing on.
Dave Vellante
>> Chris, these are your peeps, and so you've seen the evolving role of the data engineer. Do you have any thoughts on this topic?
Rebecca Knight
>> Because it's a real cultural shift within an organization.>> Yeah, sure. Well, I don't know, I'll tie this to announcements, but this is a lot of why we've focused a lot of our energy going forward on CoCo and CoWork. I think it actually fits very well into the way you were just describing this. CoCo is a set of tools to help the people who are building these data products, building the data pipelines, really getting the data ready to be consumed. And then CoWork is the platform for the people who are consuming the data. So I really think of it as there's three roles as this sort of evolves. I actually do think the names and definitions of these roles will evolve. They won't disappear and completely change, but they will evolve and look different, and we may call them different things down the line. But there are other people who are making sure that the data is in the right shape, is usable, is trustable, the right semantic models are in place. The data is set up well to be consumed by AI or by a person. And then you have the people who are going to take that data and are building the tools and agents and capabilities that your business users, the third group that you care about, are going to be able to come in and consume. But I think it will shift because the business users are no longer going to ask someone to write a query for them. Instead, they're going to ask an agent, and the data scientists will evolve to be getting the agents ready. And then you have the team that still needs to make sure the data is ready for them, but they're going to be orchestrating and managing teams of agents and other things to actually manage those pipelines. But they still need to be thinking, as you said, about what are the right pipelines? What do they cost? Do I have the right governance in place? You still need human judgment in all this.
Rebecca Knight
>> Absolutely.
Dave Vellante
>> If I go back to the Gartner variety, I'll go back there for a moment, you got structured data, unstructured data, you got streaming, you guys have announced some new Kafka capabilities, et cetera. So you've got to... I mean, the data engineer, the quality engineer, they spent a lot of time wrangling and reconciling all of that. I know you can, but specifically, how are you using R&D and product and technology to attack that problem with cross-lakehouse capabilities, et cetera?
Dave Vellante
>> So I actually think a lot of the problems that people run into coming from having different silos and different pipelines and different places that they're building and managing data. So you'll end up with multiple different teams who have each defined a certain metric in slightly different ways. That's where a lot of the wrangling and a lot of having to reconcile different data sources comes from. Generally speaking, if you go all the way back to a source system, in most cases there is ultimately a source of truth. There is a true fact that is happening somewhere in a system. And if not, if your data is entered incorrectly in Salesforce or wherever, you should fix it in the source system. Then what you need to do is have a single place that you're defining the downstream metrics and definitions that you need. So if you want to ask a question about what do we know about this Dasher, as an example, there should be a single definition of that, and that means that the place that owns that definition needs to have all of the data about that Dasher. It needs to have access to the unstructured data, the structured data, the streaming data, everything funneling into that one place so that you can have a single source of truth on what do you know about that particular fact.>> And then single source of truth is so important actually, because any use case that we are talking about in the infrastructure should talk to the same exact asset. You cannot hallucinate about it. You can go into multiple layers, go back down to the stack and get that data out, but that'll be confusing. Everybody will come back out with the wrong and incorrect information about it. Now, if you have one on think about that Dasher and have a proper information, then agentic workflow, ML workflows, or everybody should come back to the exact same system of record and make the same exact decision, so that the human and machines are taking the same exact decisions.
Dave Vellante
>> So when your agent encounters conflicting data, does the system say, "Okay, you got conflicting data, is it A or B?" and you reconcile it at that point, bring the human into the loop? First part of the question. And the second part of the question is, does the agent learn from the reasoning traces of the human and the judgment that it made so that the next time it encounters a similar conflict, it reconciles it without the human? Lond-winded question, but that's the nirvana.>> That's the hard part actually. If you expose thousands of tables to an agentic workflow and expect it to make the right call for it, it's going to make agentic workflow come back with a response and make it very confidently. Because on a keyword, it picked up a table, thought that that's the perfect information or perfect source for it, and came back with a hallucinated result and very confidently say that, "Okay, this is the number." But the point is actually about exposing the right assets, the right quality guardrails, light context and semantics around it through agentic workflows so that it understand what's the right data assets to use. Now, how do you do that actually? Which is the second part of your question.
Dave Vellante
>> Yes, right.>> How does a human analyst who is used to querying a data in a particular way for deeper analytics versus for executive reporting, so to speak, how does the agent learn that? So we are creating learning loops for the agents to understand what are the right assets, what are the context behind it, how to use them, and for in both situation, which table has to be invoked versus another. So the operator who's looking for a very detailed information gets to a lower level fact table, versus an executive who's looking for an highly aggregated data, goes to that aggregated table which has the very well-defined metrics for it. So if you don't teach what is what, then agent will come back with an incorrect response.
Dave Vellante
>> This is next-level stuff. I mean, this is relatively recent that you're actually able to do this. How has it impacted your business?
Rebecca Knight
>> People are loving it, actually. We are looking at the productivity gain. People are excited about it. Actually, what I'm seeing right now is, people are saying, "Okay, I can do all this stuff by myself." Our executives will come back and say, "I want to resolve the Dasher problem directly myself because I got this email on that." So that's super exciting to see them do it.
Rebecca Knight
>> Chris, when you're hearing VJ, I mean, obviously you work closely with VJ and have great rapport, but this is one of the largest logistics companies in the world. What would you say are some of the best practices that you are learning that you are also bringing to other customers about the success that DoorDash has had?>> It's a lot of what VJ's been talking about. I think what we're seeing repeatedly is, the companies who are able to build the best agents and deploy them are doing that because they've been very thoughtful about what data products those agents should be depending on, and those data products are well-defined. Companies that have thousands of tables in a gold layer and they try to put agents on top, it's actually very difficult because they run into points like what Dave was asking about, which is the right answer, it's not clear. Where you're much more prescriptive about these are the correct facts about Dashers, these are the correct facts about consumers, these are the correct facts about merchants, the agent is able to then reason much more effectively about it. So this is part of the... What we've taken from them is having the data engineering team elevate to be defining those facts and defining that. And facts, not really in the sense of like a fact or dim table, but really, what is the source of truth about these important pieces of information we have? Those are the ones who are able to build and deploy agents incredibly quickly. So we spend a lot of our time using the examples of customers like DoorDash who have been able to do this, and then trying to help the rest of our customer base get to that point because there is work and investment that you have to do to get to that place.
Dave Vellante
>> I want to ask you, VJ. We were in the analyst and journalist session with Sridhar. We asked them an unanswerable question. It was like, you've got the LLM vendors here, you got the SaaS vendors, they're all trying to own the system of intelligence. Obviously, Snowflake is heading in that direction. How do you see it playing out? And Sridhar said, "Look, I don't know the answer to that. Whoever wins is going to look back and rewrite history and say, 'We knew this was going to happen that way.'" It was very prescient. But what he did say is that, "Here's what I know: Product market fit is a life force, and so we need to innovate. That's how we're going to compete." So I thought that was a great answer. You were facing similar sort of competitive realities. People see what DoorDash does. Amazon is trying to get drones immediately, and you got Uber Eats out there, you got Walmart obviously competing, et cetera, et cetera, et cetera. How do you think about... Because you can't forecast the future, you don't know if drones and robots are going to be there, you obviously can take advantage of it well. So what's your mindset, your first principle on how you serve customers, Dashers, and your merchants?>> Our motto is one simple thing: Forget about the competition and think about making the product better for the three of them, for the consumer, for Dasher, and the merchant. So that's about innovation, thinking about where our product is off, and building it the right way. That has been our... I don't think we think about, as much as other companies do, about what is a competitor doing. Instead, we think about how can we make our product better, thinking about what capabilities we should build on. That's how we think about it.
Dave Vellante
>> No, that's the right answer. I mean, I think that's the mindset that Snowflake has based on what Sridhar said. I mean, because you can't predict what the competition is going to do. You can only sort of control what you can control.
Dave Vellante
>> It's innovation piece.>> Our philosophy is to pay very close attention to what our customers are doing, the problems they're running into, what they're trying to solve. Often a lot of, for us, the product development process is, look at what folks at the cutting edge like VJ are doing and then figure out how can we take that and turn it into a product that we can now sell to the rest of the customer base. That's been very successful for us of a way to find, what can we help them do? Let's work incredibly closely together and then let's go take that to everyone else.
Dave Vellante
>> But in the last three years, Chris, you guys have decided that we're not just going to be a data feed into these other intelligence systems, we're actually going to participate and add value to those intelligent systems. So that's a big strategic move that you guys->> I will say it came out of a similar thing where we saw these models come up, and we sat down with a number of customers, and they were saying, "Hey, we want to be able to use these not just with documents we have, but with all of the data that lives inside Snowflake. And we want to be able to do that in a high-trust, highly-governed way." And we tried to do that as an integration, and it didn't work very well. Part of that was because it didn't have the right context, it wasn't running within the context of the user. Companies do a lot of work around which specific rows can be accessed by which specific people for which specific purpose, and it was very hard to get that out to an LOM. So we found that the only way that we could really provide that same governance and high-quality answers to questions was to basically bring the models to the data rather than the other way around.
Dave Vellante
>> So many opportunities that will be created as a result.
Rebecca Knight
>> Indeed. It's exciting times. Well, VJ and Chris, thank you both so much for coming on the show. Really interesting conversation.
Dave Vellante
>> Thank you.>> Thank you so much. Really appreciate it.>> Thank you both.
Rebecca Knight
>> I'm Rebecca Knight for Dave Vellante. Stay tuned for more of day two of theCUBE's live coverage of the Snowflake Summit. You're watching theCUBE, the leader in enterprise tech news and analysis.