1 00:00:00,240 --> 00:00:03,360 Something breaks and then you have to jump out of the bed in the middle 2 00:00:03,520 --> 00:00:05,880 of the night to get those things fixed. 3 00:00:16,800 --> 00:00:20,080 Hello and welcome back to Data Driven, the podcast where we explore the emerging 4 00:00:20,640 --> 00:00:24,400 industry that is AI, data science, and of course, none of 5 00:00:24,400 --> 00:00:28,250 it's all possible without data engineering. Now, unfortunately, 6 00:00:28,250 --> 00:00:31,850 my favorite, most favorite data engineer in the world can't make it here 7 00:00:32,170 --> 00:00:35,850 today. And I am actually enjoying the sunny but not too 8 00:00:36,090 --> 00:00:39,530 ridiculously hot sunny day here in the suburbs of Baltimore, 9 00:00:39,930 --> 00:00:43,770 Maryland. But I am excited here because I have Pranesh Patel, 10 00:00:43,850 --> 00:00:47,690 who is the co-founder and CEO of a company that 11 00:00:47,770 --> 00:00:51,530 does— makes data engineering a lot more palatable, it sounds like, Ultima 12 00:00:51,690 --> 00:00:54,940 AI. Welcome to the show. Hey, um, 13 00:00:55,580 --> 00:00:59,020 Frank, thanks for having me here. Super excited to chat with you 14 00:01:00,200 --> 00:01:03,340 today. Yeah, so, so tell me about your company, Ultimate AI. Is, uh, 15 00:01:03,980 --> 00:01:06,620 tell me about what it is and what led you to make it. 16 00:01:08,620 --> 00:01:12,380 So Ultimate AI, we are a startup based out of the San Francisco Bay Area. 17 00:01:13,420 --> 00:01:16,860 What we do is we use AI to 18 00:01:17,740 --> 00:01:21,420 automate a bunch of data engineering tasks for the teams out there 19 00:01:21,660 --> 00:01:24,620 who do data work, right? The tasks like building 20 00:01:25,660 --> 00:01:29,020 ELT pipeline, extract and loading of the data, transforming 21 00:01:29,180 --> 00:01:32,940 data, and not just building those pipelines but maintaining 22 00:01:33,260 --> 00:01:37,100 those as well. It's always nightmares. Something breaks and then you have to 23 00:01:37,180 --> 00:01:39,740 jump out of the bed in the middle of the night to get those things 24 00:01:39,980 --> 00:01:43,660 fixed. So building and maintaining data pipelines or 25 00:01:44,140 --> 00:01:47,780 managing your data infrastructure Like in the modern data 26 00:01:48,020 --> 00:01:51,620 stack today, we have Snowflake, Databricks, BigQuery, those things, right? 27 00:01:52,180 --> 00:01:55,940 We need to spend a bunch of effort to manage that infrastructure as well as 28 00:01:56,420 --> 00:02:00,260 optimize that infrastructure, like making sure you have turned on the right knobs on the 29 00:02:00,580 --> 00:02:04,340 infrastructure side as well as you have optimized those queries and pipelines, 30 00:02:04,580 --> 00:02:08,420 etc. So we work on that as well. The whole idea is how we 31 00:02:08,500 --> 00:02:12,320 can use the agents to automate a bunch of that work. so 32 00:02:12,560 --> 00:02:16,320 we can scale our teams even further. I'm glad 33 00:02:16,400 --> 00:02:20,080 you mentioned— I'm sorry, go ahead. No, you asked 34 00:02:20,320 --> 00:02:23,920 me that, hey, how did it all start? So me and my co-founder, 35 00:02:24,720 --> 00:02:28,480 we have worked in B2B enterprise space for a long time. You must have seen 36 00:02:28,640 --> 00:02:32,320 all those data teams every company has. There's so much work involved. There is a 37 00:02:32,400 --> 00:02:36,000 long backlog. We went through the same experiences, and we started 38 00:02:36,320 --> 00:02:39,920 thinking there must be a better solution so that we can handle that 39 00:02:40,080 --> 00:02:43,680 backlog. At the same time, a few years ago, this whole AI wave started 40 00:02:43,920 --> 00:02:47,520 coming in. We jumped in headfirst and started building 41 00:02:47,760 --> 00:02:51,280 agents first to sort of automate simple tasks like writing 42 00:02:52,320 --> 00:02:56,080 documentation or writing some data quality tests. And through that, 43 00:02:56,320 --> 00:03:00,160 the tool and the product evolved from there. And now 44 00:03:00,240 --> 00:03:04,080 we automate bunch of things that are day-to-day things for 45 00:03:04,160 --> 00:03:08,010 data engineering folks. That's a good point. You brought up a good point 46 00:03:08,170 --> 00:03:11,770 because you mentioned Snowflake, you mentioned Databricks, you mentioned all of these 47 00:03:12,170 --> 00:03:15,930 technology tools that the stack used to be a lot simpler in data, 48 00:03:16,250 --> 00:03:19,050 in the data space, right? You, you know, you were either an Oracle shop 49 00:03:20,090 --> 00:03:23,850 or a SQL Server shop and you had, you know, SQL 50 00:03:24,010 --> 00:03:27,610 Server tooling, which those I'm way more familiar with, but also the 51 00:03:27,690 --> 00:03:31,210 Oracle kind of stack too, right? So it went from being kind of a 52 00:03:32,820 --> 00:03:36,260 The tools may have been limited, yes, but the tools also— 53 00:03:37,540 --> 00:03:40,500 there only wasn't really that much of them, right? You picked one side 54 00:03:41,620 --> 00:03:45,380 and stuck with it, right, for most organizations. And then we 55 00:03:45,460 --> 00:03:48,740 had things like Hadoop and Pig and Hive and all of those things kind of 56 00:03:48,740 --> 00:03:52,580 come out. And now this many decade and a half or so 57 00:03:53,060 --> 00:03:56,820 or more, you can tell by my gray hair, that now we 58 00:03:56,980 --> 00:04:00,280 have just an unlimited assortment, it seems, of these 59 00:04:00,440 --> 00:04:03,880 tools, right? And each one of them has their own quirks, their own settings, as 60 00:04:03,960 --> 00:04:06,920 you said, the knobs and dials. Is that something that your 61 00:04:08,120 --> 00:04:11,160 solution offers a solution to? Yeah, 62 00:04:11,160 --> 00:04:14,760 absolutely. Because as you mentioned, the use 63 00:04:15,000 --> 00:04:18,280 cases exploded, and with that, the tool stack exploded as 64 00:04:18,440 --> 00:04:22,280 well. You can't be exploiting 35 different tools to make sure 65 00:04:22,760 --> 00:04:25,800 everything works perfectly and keep your tabs on everything. 66 00:04:26,580 --> 00:04:30,180 So we help tremendously with that, like all those configurations in your 67 00:04:30,660 --> 00:04:34,500 infrastructure as well as building out things. You can't be expert in so many tools. 68 00:04:34,740 --> 00:04:37,860 And on the other side with AI, what has happened is 69 00:04:38,580 --> 00:04:42,420 even these tools as a software are evolving so fast. In a 70 00:04:42,500 --> 00:04:46,260 year, there are like 100 features come out. How can you keep tab on what's 71 00:04:46,260 --> 00:04:49,940 the latest and greatest and what actually fits in my environment and to my use 72 00:04:50,100 --> 00:04:53,860 cases? And, you know, AI for Rescue there, it can 73 00:04:53,940 --> 00:04:57,740 do a bunch of those things for you. No, absolutely. And 74 00:04:58,220 --> 00:05:01,740 I hadn't really thought about that angle of you have to be an expert. 75 00:05:02,140 --> 00:05:05,260 It's like what happened in the software development world 76 00:05:06,300 --> 00:05:10,140 where you had the notion of a full-stack developer where suddenly, 77 00:05:10,940 --> 00:05:14,540 you know, it went from you would be a Visual Basic developer, right? 78 00:05:14,940 --> 00:05:18,780 Or a developer of, you know, web developer. But 79 00:05:18,860 --> 00:05:21,740 then it became, no, you had to be a full-stack developer. You had to know 80 00:05:21,820 --> 00:05:24,140 the data side. You had to know the CSS. You had to know the HTML 81 00:05:24,220 --> 00:05:27,680 and the JavaScript, right? Yeah. Data, I think, has followed a very similar trajectory 82 00:05:28,000 --> 00:05:31,680 in that regard of no one person can do it all. At least 83 00:05:33,040 --> 00:05:36,640 no one person can do it all without the assistance of some kind of 84 00:05:36,640 --> 00:05:40,480 AI. And is that what your product enables? Like 85 00:05:40,640 --> 00:05:44,240 somebody to like basically be a— in the DC area, 86 00:05:44,480 --> 00:05:47,840 they love the term force multiplier. Is that kind of like what 87 00:05:48,320 --> 00:05:51,930 your tool does? Exactly. Exactly. Because 88 00:05:52,650 --> 00:05:56,250 as I think things exploded, use cases exploded, and even 89 00:05:56,490 --> 00:06:00,170 the use of AI has exploded overall, your data 90 00:06:00,410 --> 00:06:04,250 actually fuels all of it. And what's happening is the data projects 91 00:06:04,570 --> 00:06:08,410 are growing exponentially, but our teams are not growing in that size. And 92 00:06:08,410 --> 00:06:12,250 that's why there is a huge backlog that's happening. So the idea is then, 93 00:06:12,730 --> 00:06:16,570 in your words, how we can have some sort of force multiplier that 94 00:06:16,730 --> 00:06:20,570 can do a bunch of things for me automatically. And I focus on some of 95 00:06:20,570 --> 00:06:24,410 the more complex things or something where AI needs more help 96 00:06:24,570 --> 00:06:28,170 and guidance. So think of this as these bunch of agents 97 00:06:28,490 --> 00:06:32,010 are my minions. They're getting the work done. I have like 30 of those and 98 00:06:32,090 --> 00:06:35,850 I'm giving them directions, course correcting, telling them what to do. And 99 00:06:35,930 --> 00:06:39,770 that's how basically the teams are scaling themselves today. But there are so many challenges 100 00:06:40,170 --> 00:06:43,770 in doing that as well, in that whole process. And our 101 00:06:44,010 --> 00:06:47,860 goal is how we can streamline and make that whole process smooth. So you have 102 00:06:48,020 --> 00:06:51,580 teams of these minions which are getting a lot of work done for you. 103 00:06:52,100 --> 00:06:55,780 Yeah, and I think you brought up a very real pain 104 00:06:55,940 --> 00:06:59,380 point, right? They're not hiring tens of people in the data engineering 105 00:06:59,700 --> 00:07:03,460 teams anymore, right? Data engineers are expected to do, to all 106 00:07:03,540 --> 00:07:06,740 be 10x engineers, right? And so 107 00:07:07,780 --> 00:07:11,380 I like the idea of you calling these AI agents minions, one, 'cause 108 00:07:11,620 --> 00:07:14,920 I think the movies are really cute and I have small kids. But 109 00:07:16,040 --> 00:07:19,240 how do you— what's the governance look like on that? Is that something that you— 110 00:07:19,560 --> 00:07:23,080 how did you address that problem? 'Cause I'm sure that comes up quite a bit. 111 00:07:24,120 --> 00:07:27,880 No, I think you touched on very important area, right? Because as a 112 00:07:28,040 --> 00:07:31,640 human, we usually have general sense of understanding and 113 00:07:31,880 --> 00:07:35,720 what needs to be done and what shouldn't be done. But for agents, 114 00:07:35,960 --> 00:07:39,480 it's a piece of code, it's a machine, right? A lot of times it 115 00:07:39,560 --> 00:07:43,100 doesn't have that, compass of common sense. So for example, 116 00:07:43,180 --> 00:07:46,780 that's why you might have seen situations where agent ran a query and that cost 117 00:07:46,940 --> 00:07:50,540 you thousands of dollars. Oh, and there's so many stories of agent 118 00:07:50,860 --> 00:07:54,700 deleted my sensitive data or replicated it. So whether it's sensitive 119 00:07:55,020 --> 00:07:58,460 data, access controls, cost guardrails, 120 00:07:59,260 --> 00:08:02,380 all of those are very important factors around which 121 00:08:03,180 --> 00:08:06,870 we need to develop a solution. Now we just talked about Lot 122 00:08:06,950 --> 00:08:10,550 of tools in the data stack. Every tool has their access layer, RBAC 123 00:08:10,710 --> 00:08:14,550 layer. Now, how are we going to do this, right? So what we have been 124 00:08:14,550 --> 00:08:18,310 doing is, and that's the reason we launch the, our product as open 125 00:08:18,550 --> 00:08:21,910 source project, we call it Ultimate Core. There is a governance 126 00:08:22,390 --> 00:08:26,230 layer that's built in, which allows people to define these guardrails, rules, 127 00:08:26,550 --> 00:08:29,910 and permissions. Lot of guardrails come inbuilt. 128 00:08:31,030 --> 00:08:34,630 And the beauty of that is this is a layer that sits on top of 129 00:08:34,790 --> 00:08:38,529 these bunch of different tools that you are using in your data stack. and gives 130 00:08:38,609 --> 00:08:42,049 us the common ground around governance. So for example, 131 00:08:42,129 --> 00:08:45,569 cost, it makes sure the agent doesn't spend more money than this 132 00:08:46,209 --> 00:08:49,569 for your specific task. Or for agent 133 00:08:50,049 --> 00:08:53,809 itself, you can create a separate view, separate tables so that they don't 134 00:08:54,049 --> 00:08:57,809 inherit like service user permissions or user permissions directly. Because as 135 00:08:57,889 --> 00:09:01,729 a user, I might have so many permissions, but I don't want my agent to 136 00:09:01,809 --> 00:09:05,329 have the same permissions and use my same credentials, et cetera. 137 00:09:05,850 --> 00:09:08,170 So that part also we have solved pretty well. 138 00:09:09,850 --> 00:09:13,450 Yeah, I think, I think calling them minions works out pretty well too, because, you 139 00:09:13,530 --> 00:09:17,370 know, the minions always— there's a whole sequence in the first movie about the dart 140 00:09:17,530 --> 00:09:20,570 gun. Parents, if you know, you know. But 141 00:09:21,370 --> 00:09:23,530 the minions misheard it and built something else. But 142 00:09:25,370 --> 00:09:29,130 no, I think you're right. Making sure these things have guardrails on them, 143 00:09:29,210 --> 00:09:33,050 I think, is— people are going to learn very quickly what happens if you don't 144 00:09:33,050 --> 00:09:36,220 have guardrails, right? millions of dollars a query, or etc., etc. 145 00:09:37,180 --> 00:09:40,940 Actually, recently in the news, and they haven't disclosed what was the core— what 146 00:09:41,020 --> 00:09:44,860 was the core problem, but apparently AWS had— 147 00:09:45,260 --> 00:09:48,300 was giving out bills that were like orders of magnitude higher. 148 00:09:49,580 --> 00:09:53,420 And I can only— again, I have no inside information, 149 00:09:53,580 --> 00:09:57,100 but I can only imagine that there was probably some kind of rogue AI doing 150 00:09:57,260 --> 00:10:01,020 a math error or doing something crazy like that. What do you think the 151 00:10:01,100 --> 00:10:04,480 future of these— this tool space is going to be? Do you think You know, 152 00:10:05,200 --> 00:10:08,320 there'll be more MCP servers, right? What do you think, 153 00:10:08,560 --> 00:10:12,400 um, what do you think that's going to look like, the ecosystem, 154 00:10:12,400 --> 00:10:16,160 the data ecosystem? So how I 155 00:10:16,960 --> 00:10:20,480 see it is based on what we want to do, the tools 156 00:10:20,800 --> 00:10:23,920 get built, right? And as we sort of expand our 157 00:10:24,320 --> 00:10:28,160 horizons to do more and more things, the tools get expanded also, right? 158 00:10:28,320 --> 00:10:32,030 For example, I think MCP server is a great example. At one point 159 00:10:32,190 --> 00:10:35,630 we realized, hey, we need to sort of feed in all this 160 00:10:36,030 --> 00:10:39,630 information to MCP server, uh, to AI agents, right? And then 161 00:10:39,710 --> 00:10:43,150 MCP servers were built as a solution to it. Now 162 00:10:43,710 --> 00:10:47,470 we are hitting the limits of MCP servers themselves where, 163 00:10:47,790 --> 00:10:51,630 you know, okay, governance— anybody can install any MCP server and I have 164 00:10:51,710 --> 00:10:55,470 no control over who uses which MCP server in the organization. Right. That's 165 00:10:56,190 --> 00:10:59,980 happening. Second is The MCB servers are putting out so 166 00:11:00,140 --> 00:11:03,900 much output that my tokens are getting consumed like candy, 167 00:11:03,980 --> 00:11:07,580 right? So then how do I control it where MCB tool outputs are 168 00:11:07,660 --> 00:11:11,020 limited? And we have built some functionality around that, around context comparison 169 00:11:12,140 --> 00:11:15,180 and tool output curtails, etc., especially for data engineering 170 00:11:15,340 --> 00:11:19,020 tasks. Our third big thing is MCB server is like 171 00:11:19,340 --> 00:11:22,380 API, is going to just pull the data from the other system and bring it 172 00:11:22,460 --> 00:11:25,830 to you. But it is not going to do any intelligent filtering 173 00:11:26,470 --> 00:11:29,750 or connecting those dots together. And then now people are looking at 174 00:11:30,070 --> 00:11:33,430 semantic layers and those things, how I can do it also. So now we are 175 00:11:33,510 --> 00:11:37,270 running into the limitations of NCP servers themselves, and then people are trying 176 00:11:37,350 --> 00:11:40,870 to figure out, hey, what's the next thing we need to do to solve this 177 00:11:41,110 --> 00:11:44,790 problem really well? And then people are talking about context graphs 178 00:11:44,950 --> 00:11:48,470 or context store where all of your information is already there, 179 00:11:48,790 --> 00:11:52,390 curated, filtered, smartly arranged, And then you use MCP 180 00:11:52,630 --> 00:11:56,230 server to pull only right information instead of directly interfacing with the 181 00:11:56,790 --> 00:12:00,550 tool, like for example Salesforce or HubSpot, and just dumping everything into your 182 00:12:00,630 --> 00:12:04,310 AI agent as well. So how I see it is, I think 183 00:12:04,550 --> 00:12:08,390 agents are going to become more and more autonomous, more and more ambient 184 00:12:08,630 --> 00:12:12,230 as well, and the information we are going to feed them 185 00:12:12,470 --> 00:12:16,230 is going to become much more curated as well. And there are so 186 00:12:16,310 --> 00:12:20,070 many facets to this. There is a right information feeding angle around 187 00:12:20,390 --> 00:12:23,970 context. There is a governance angle to it, and now. And now I think you 188 00:12:24,050 --> 00:12:27,730 just touched on it. The one big angle that's coming into the play is cost 189 00:12:27,890 --> 00:12:31,730 as well. There, there have been so many stories coming out where people spend 190 00:12:31,890 --> 00:12:35,650 their entire year's worth of budget in like 3 months. I know some stories 191 00:12:35,890 --> 00:12:39,730 which are not public where people's usage like 30x'd in 192 00:12:39,890 --> 00:12:43,490 6 months and now they're like, oh my God, my AI bill is actually same 193 00:12:43,650 --> 00:12:46,610 as my cloud bill now. I never planned for this. 194 00:12:47,490 --> 00:12:51,090 And what is exactly the ROI that people are trying to measure also? 195 00:12:51,330 --> 00:12:54,550 Yeah. So all these questions are coming up, and as the questions and use cases 196 00:12:54,870 --> 00:12:58,070 come up, I believe we'll have more tooling and better solutions. 197 00:13:00,230 --> 00:13:04,070 Is anything— is there anything in particular that your product addresses 198 00:13:05,350 --> 00:13:08,950 to any of these problems? Yeah, so we talked about the governance 199 00:13:09,430 --> 00:13:13,110 piece of it. On the cost side as well, what we started doing 200 00:13:13,270 --> 00:13:17,110 is one of the features we have in the product is context compaction. 201 00:13:17,350 --> 00:13:21,000 So I was talking about MCP tools. Dumping a lot of data, like 202 00:13:21,240 --> 00:13:24,520 especially data-related MCP tools. So what we do is 203 00:13:25,080 --> 00:13:28,760 we curtail the output correctly. We know all these MCP 204 00:13:28,920 --> 00:13:32,440 servers, etc., so that you are not spending too much of a token cost. And 205 00:13:32,520 --> 00:13:36,360 not just the cost, right? What they call is a context rot. If you 206 00:13:36,520 --> 00:13:40,280 dump in too much information, LLM has too many directions to go in. So we 207 00:13:40,360 --> 00:13:44,120 curtail and do that context compaction. Second part we 208 00:13:44,200 --> 00:13:48,010 have introduced is Specifically for data tasks, we create 209 00:13:48,410 --> 00:13:52,170 memories automatically. So not every time agent is starting from scratch, 210 00:13:52,410 --> 00:13:56,250 and it's very curated for data-related tasks. So in that way, 211 00:13:56,410 --> 00:13:59,770 next time when the agent does the same task, it can tap into those memories 212 00:14:00,170 --> 00:14:03,930 that are shared across the organization and get the task done in 213 00:14:04,330 --> 00:14:08,090 very less number of steps. That's an interesting 214 00:14:08,250 --> 00:14:11,540 point because I noticed that, that was one of the When I started 215 00:14:12,980 --> 00:14:16,420 poking around my OpenClaw instance, right, inside of there, there's the 216 00:14:16,580 --> 00:14:20,100 soul, but there's also kind of this memory type of notion and managing 217 00:14:20,500 --> 00:14:24,100 that memory so the context doesn't have to start from 218 00:14:24,340 --> 00:14:28,180 zero every time. And does that really save on tokens? Like, what's 219 00:14:28,180 --> 00:14:31,060 the rough order of magnitude in terms of what the token save is? 220 00:14:32,180 --> 00:14:35,940 No, the memory will save you tremendously because, for example, 221 00:14:36,100 --> 00:14:39,530 let's take an example of data pipelines. Right? So if your data 222 00:14:39,690 --> 00:14:43,530 pipelines are failing, usually there is a pattern, same issues you're going 223 00:14:43,690 --> 00:14:47,290 to see again and again. Hey, data hasn't landed, that's why this pipeline particularly 224 00:14:47,690 --> 00:14:51,450 fails. Now if that thing gets stored in the memory, your agent 225 00:14:51,690 --> 00:14:54,810 is going to check the first thing is has data already 226 00:14:54,970 --> 00:14:58,730 landed? That will save you other 10 other things 227 00:14:59,370 --> 00:15:03,050 that agent would normally try out before coming to that, right? Boom, your 228 00:15:03,290 --> 00:15:06,900 multiple workflows are saved. Your token cost saved, your time is saved 229 00:15:07,140 --> 00:15:10,660 also. We fixed it very quickly, right? So it helps us 230 00:15:11,060 --> 00:15:14,740 tremendously in that way. And the beauty of that is even the memory 231 00:15:15,140 --> 00:15:18,900 layer cannot be very generic, right? You need to understand 232 00:15:19,780 --> 00:15:23,620 as tasks are happening, what kind of memories are important. It needs 233 00:15:23,780 --> 00:15:27,380 to be curated. There needs to be, say for example, somebody's writing 234 00:15:28,100 --> 00:15:31,930 agents and harness for finance, there needs to be a finance-specific memory. that 235 00:15:32,170 --> 00:15:35,930 will remember finance-related things. Similarly, for data work, what we 236 00:15:36,010 --> 00:15:39,770 have created is data agent-specific memories, which works amazingly 237 00:15:39,770 --> 00:15:43,290 well. And since we are talking about memory, right, memory 238 00:15:43,610 --> 00:15:47,130 is not just limited to one session. So for example, what I'm trying to say 239 00:15:47,130 --> 00:15:50,970 is, now if you have data pipelines, probably you have a team of 240 00:15:51,370 --> 00:15:55,210 engineers maintaining that data pipeline. So for example, 241 00:15:55,290 --> 00:15:58,900 I fixed this data pipeline today, for that data not landing 242 00:15:59,140 --> 00:16:02,660 issue, and it gets saved in my memory. Of course, next time 243 00:16:02,980 --> 00:16:06,340 in the regular systems, I try to fix it. Next time it will read from 244 00:16:06,420 --> 00:16:10,260 the memory stored in my, say, code editor or Cloud Code or something like that, 245 00:16:10,260 --> 00:16:13,780 and I can draw to it. But what about somebody else on my team? They 246 00:16:13,780 --> 00:16:17,220 are also going to work on that pipeline, and they might be on on-call. That's 247 00:16:17,220 --> 00:16:20,740 when the pipeline failed. So the memory layer that we have 248 00:16:20,900 --> 00:16:24,260 built, it's a layer that gets shared across the teams and 249 00:16:24,740 --> 00:16:28,530 organization as well. So then it's the— I call this a 250 00:16:28,610 --> 00:16:32,450 hive-like mind in which you are storing the memory, and anybody can 251 00:16:32,610 --> 00:16:36,370 come in and use that hive-like mind and sort of use 252 00:16:36,530 --> 00:16:39,730 that knowledge to do things better. And the big benefit of this is 253 00:16:40,530 --> 00:16:44,130 even for newer people joining your team, right, they don't have to start from scratch. 254 00:16:44,770 --> 00:16:48,610 All this tribal knowledge is stored in that hive mind, which 255 00:16:48,690 --> 00:16:52,540 we call as a memory layer. I like that because then 256 00:16:53,660 --> 00:16:57,500 your first few rounds of learning it, right, you're 257 00:16:57,580 --> 00:17:01,340 literally onboarding this virtual employee, it sounds like, right? It sounds somewhere between a minion 258 00:17:02,220 --> 00:17:05,900 and a virtual employee that can kind of capture that tribal knowledge, 259 00:17:06,140 --> 00:17:09,740 capture that institutional kind of wisdom. Yeah. And so 260 00:17:09,740 --> 00:17:13,500 it's not so much you're spending tokens, you're kind of investing tokens, right, 261 00:17:13,660 --> 00:17:16,780 for the future and training this virtual employee. And 262 00:17:17,180 --> 00:17:19,980 presumably you'll get that, you'll see dividends later on. 263 00:17:20,790 --> 00:17:24,310 Exactly. And this system works with agentic 264 00:17:24,790 --> 00:17:28,630 frameworks that are out there already. People use, say, Claude Code, GitHub 265 00:17:28,790 --> 00:17:32,150 Copilot, Cursor. The system works with those 266 00:17:32,470 --> 00:17:35,750 already. We don't build LLMs, we don't build— 267 00:17:36,550 --> 00:17:40,230 give agentic frameworks, but we build this harness which makes 268 00:17:40,710 --> 00:17:44,470 these tools extremely suitable or powerful when it 269 00:17:44,470 --> 00:17:45,830 comes to data engineering work. 270 00:17:48,280 --> 00:17:51,960 Interesting. Can— does it learn on its own or 271 00:17:52,200 --> 00:17:55,720 can you edit those memories? Right. So like, what if— what, 272 00:17:55,880 --> 00:17:59,320 let's just say we have a, somebody makes a mistake, right? You obviously 273 00:17:59,720 --> 00:18:03,320 wanna mark that and kind of remove that from the memory. Is that, is that 274 00:18:03,640 --> 00:18:07,400 like, obviously you probably edit, it probably adds its own. And, 275 00:18:07,480 --> 00:18:10,440 and the reason why I mentioned this is because when I dove into my 276 00:18:11,650 --> 00:18:15,490 my OpenCLAWS memory file. I thought it was funny what it read about me. 277 00:18:15,650 --> 00:18:18,930 So if people are watching this, you'll see I'm kind of like winking my eyes. 278 00:18:19,010 --> 00:18:22,850 Apparently allergies are really bad today, which I did not factor that in when 279 00:18:23,010 --> 00:18:26,770 sitting outside. And one of the things it learned about me was I 280 00:18:26,850 --> 00:18:30,210 always ask about the pollen levels for the day, which I think is kind of 281 00:18:30,210 --> 00:18:33,170 funny. It said that, you know, I ask about the weather, I ask about stocks, 282 00:18:33,730 --> 00:18:37,010 and I ask about, you know, AI innovations and pollen report, 283 00:18:37,250 --> 00:18:40,840 right? Does this— but obviously I can go in, I can open up a terminal 284 00:18:41,080 --> 00:18:44,920 and edit it myself. But how does your solution— is it built into 285 00:18:44,920 --> 00:18:48,200 the UI or is it like just a config file? So 286 00:18:49,320 --> 00:18:52,920 as you start using the solution, you install, it will start creating 287 00:18:53,320 --> 00:18:57,160 memories automatically. You can, of course, in your prompt give a 288 00:18:57,240 --> 00:19:01,000 specific instruction. As you give the instructions, those get saved also. But you 289 00:19:01,080 --> 00:19:04,600 can say specifically create a memory also. And through our 290 00:19:04,760 --> 00:19:08,580 MCP, it will get that memory created as well. Now the harder 291 00:19:08,820 --> 00:19:12,340 part usually is what if there are bad memories and you want to change memories 292 00:19:12,820 --> 00:19:16,340 or you want to erase those out? The good news is it comes with that 293 00:19:16,980 --> 00:19:20,580 flash stick that I think I remember it from the movie Men in 294 00:19:20,660 --> 00:19:24,500 Black, right? It comes with that. So you can go in, delete your 295 00:19:24,660 --> 00:19:28,500 memories because as I was telling you, we store these memories in a SaaS. So 296 00:19:28,580 --> 00:19:32,420 in that way, they're shared with your team members, they're shared with the rest 297 00:19:32,500 --> 00:19:36,020 of the organization, and there is a granular control you can do who they get 298 00:19:36,180 --> 00:19:39,270 shared with. But at the same time, if you want to update 299 00:19:39,430 --> 00:19:42,950 it, you can go to the UI and get those updated or 300 00:19:43,270 --> 00:19:47,030 deleted as well. Oh, interesting. So you 301 00:19:47,110 --> 00:19:50,230 have that neuralyzer built in. I think that's what the thing is called in Men 302 00:19:50,310 --> 00:19:53,910 in Black. But so you can go back 303 00:19:54,150 --> 00:19:57,510 and you can remove like, hey, everything I did today was terrible, so don't remember 304 00:19:57,670 --> 00:20:01,110 that. What about 305 00:20:01,830 --> 00:20:05,490 security, right? Obviously, There's a lot of trade, you know, 306 00:20:05,570 --> 00:20:08,690 there's a lot of sensitive information that are going to be floated around in the 307 00:20:08,770 --> 00:20:12,530 data engineering space. How does your solution, how does Ultimate AI 308 00:20:12,850 --> 00:20:16,610 kind of address that? So first and foremost, 309 00:20:16,690 --> 00:20:20,530 we don't look at the data directly. Okay, so it's a 310 00:20:20,530 --> 00:20:24,130 harness that comes in, and this harness people can install locally 311 00:20:24,450 --> 00:20:28,210 as well. So if you're some sensitive industry, maybe healthcare 312 00:20:28,530 --> 00:20:31,200 or something like that, You can use it completely 313 00:20:31,520 --> 00:20:35,040 locally, and the LLM solution that you use in the 314 00:20:35,120 --> 00:20:38,960 background, people can use their LLM solution. As I was talking about, Claude 315 00:20:39,120 --> 00:20:42,480 Core subscription or Codex subscription, they can use 316 00:20:42,640 --> 00:20:46,400 that. They can hook up their own models also. We support, say for example, 317 00:20:46,720 --> 00:20:50,480 OpenRouter, all these different models people can use. Even if they want to use 318 00:20:51,040 --> 00:20:54,720 on-premise LLM, we support that as well. On the other side, 319 00:20:54,880 --> 00:20:58,030 as a company and as a platform, We are SOC 2 certified, 320 00:20:58,830 --> 00:21:01,790 pen tested, a bunch of big Fortune 500 companies 321 00:21:02,830 --> 00:21:06,590 use our product already. So even on that side, I'm 322 00:21:06,590 --> 00:21:10,110 sure we can make people's security teams happy because we have done a bunch of 323 00:21:10,190 --> 00:21:13,790 work around it already. Oh, that's interesting. That's good. 324 00:21:14,110 --> 00:21:17,550 Because I mean, the security conversation 325 00:21:19,390 --> 00:21:22,430 comes up in AI, but I don't think it comes up often enough or early 326 00:21:22,750 --> 00:21:26,600 enough. And Obviously, I think that's going to change as more and 327 00:21:26,600 --> 00:21:30,440 more systems get deployed. And obviously, the Fortune 328 00:21:30,520 --> 00:21:34,040 500 companies obviously also take that into account as 329 00:21:34,120 --> 00:21:37,960 well. What would be your advice to people who are data 330 00:21:38,440 --> 00:21:41,560 engineers who are curious about how do I make— how do I get my own 331 00:21:41,720 --> 00:21:45,480 minions, right? Like, how do I start thinking about— let's roll 332 00:21:45,560 --> 00:21:49,240 that up. How do I start thinking about 333 00:21:49,560 --> 00:21:52,740 my job as a data engineer in terms of 334 00:21:53,700 --> 00:21:57,380 I'm managing a dozen potential virtual employees 335 00:21:57,540 --> 00:22:00,940 as opposed to doing it myself, quote unquote, the old-fashioned way. 336 00:22:01,620 --> 00:22:04,500 Yeah, yeah. No, I think that's how we should start 337 00:22:05,540 --> 00:22:08,740 thinking about it if anybody hasn't started going on that 338 00:22:08,980 --> 00:22:12,580 path. There are a bunch of tools out there, and I 339 00:22:12,820 --> 00:22:16,580 believe it's easy to get lost also because there are just so many tools and 340 00:22:16,660 --> 00:22:20,220 things that are happening. on the AI side of the things. And what I have 341 00:22:20,300 --> 00:22:24,140 seen is people try to retrofit the tools for software engineers to data engineering, 342 00:22:24,300 --> 00:22:28,060 and usually that leaves bad taste in their mouth. And sometimes 343 00:22:28,380 --> 00:22:31,900 I heard, hey, AI doesn't work. It's because you're using the wrong tools for data 344 00:22:32,220 --> 00:22:35,500 engineering most of the times. So you need to use the 345 00:22:35,660 --> 00:22:39,420 harnesses or tools that are specifically done for data engineering. 346 00:22:39,580 --> 00:22:43,420 On our side, we open source Ultimate Code Project, so definitely try it out. 347 00:22:43,500 --> 00:22:47,270 It's MIT licensed, completely open source. But in addition to that, what 348 00:22:47,350 --> 00:22:51,030 we have done is on our website, we have put this 349 00:22:51,190 --> 00:22:54,310 course called AI Data Engineer. So there is a section on our 350 00:22:55,590 --> 00:22:59,430 ultimate.ai website, AI Data Engineer, where we have put 351 00:22:59,670 --> 00:23:02,950 curated articles and we keep spending time to make sure all the updated 352 00:23:03,510 --> 00:23:07,110 information is there. So somebody might be like at day 1, somebody might 353 00:23:07,510 --> 00:23:11,270 already be at day 50. You go through that material and it will coach 354 00:23:11,510 --> 00:23:14,940 you. What are the latest tools and what are the best things you can possibly 355 00:23:15,340 --> 00:23:19,100 use to get these things done? Oh, that's really cool. So you 356 00:23:19,180 --> 00:23:22,940 have built-in training, and you did mention that your product is open source. We'll make 357 00:23:22,940 --> 00:23:26,620 sure that the link to the GitHub repo and as well as the, um, 358 00:23:27,260 --> 00:23:31,100 this onboarding, um, this kind of this ment— what would you call it, 359 00:23:31,180 --> 00:23:35,020 a mentoring tool, an onboarding tool? Like, what would you call 360 00:23:35,180 --> 00:23:38,780 that? I just call it training boot camp. So 361 00:23:39,100 --> 00:23:42,440 join it. Even if you're brand new, there is a bunch of useful 362 00:23:43,000 --> 00:23:46,840 information. Even if you have been using it for a while, you know, world 363 00:23:47,000 --> 00:23:50,840 of AI and technology changes so fast. Definitely it will help you 364 00:23:50,920 --> 00:23:53,000 keep tabs on what are the things changing as well. 365 00:23:54,760 --> 00:23:58,280 Yeah, no, I— it's a, it's a very fast-moving space 366 00:23:58,600 --> 00:24:02,440 and it, it's not gotten slower. It's— I think the pace of innovation is 367 00:24:02,520 --> 00:24:05,080 accelerating for good or for bad. What 368 00:24:07,970 --> 00:24:11,810 What would be your advice to someone? I started asking this in the question, like, 369 00:24:12,290 --> 00:24:16,050 there's a lot of computer science students that are very still 370 00:24:16,130 --> 00:24:19,970 in school and they're very worried about AI taking their jobs. I think you and 371 00:24:20,050 --> 00:24:23,810 I have gotten to the point where AI doesn't really take away your job. I 372 00:24:23,890 --> 00:24:26,850 think it changes the nature of your job. 373 00:24:27,490 --> 00:24:31,330 Yeah. But what would be first? My advice would be stick it 374 00:24:31,410 --> 00:24:34,760 through, kids. But like aside from that, what would be 375 00:24:35,160 --> 00:24:38,760 your advice to particularly, so I think a lot of people 376 00:24:39,720 --> 00:24:43,400 are getting started in data engineering and if they hear that there's AI tools 377 00:24:43,720 --> 00:24:47,320 assisting with that or doing that, they may be a little bit afraid or concerned 378 00:24:47,560 --> 00:24:51,160 about the future here. I think the future looks bright, but that's just my opinion 379 00:24:51,880 --> 00:24:55,000 as a, I wouldn't say I'm a perpetual 380 00:24:55,400 --> 00:24:58,600 optimist, but I've kind of been through that cycle of, 381 00:24:59,240 --> 00:25:03,010 oh, you know, I, AI is taking away my job, but I really kind of 382 00:25:03,090 --> 00:25:05,730 see the nature of it. It just changes the nature of my job. 383 00:25:06,930 --> 00:25:10,610 Yeah, yeah. If I can summarize this, I have heard those 384 00:25:10,850 --> 00:25:14,690 things, hey, is AI going to take my job? And I always tell 385 00:25:14,930 --> 00:25:18,530 if anybody says this to me, in my opinion, 386 00:25:18,690 --> 00:25:22,370 it's not AI that's going to take your job, but it's going to be somebody 387 00:25:22,690 --> 00:25:26,050 who can use AI is going to take your job. So if you don't upscale 388 00:25:26,530 --> 00:25:30,130 yourself, because I completely resonate with your point. Our 389 00:25:30,770 --> 00:25:34,370 jobs and roles are changing rapidly in this new world, so 390 00:25:34,530 --> 00:25:38,050 upskilling yourself for AI is very, very important. And being 391 00:25:38,450 --> 00:25:42,050 on that cutting edge of the AI and data engineering and all that data 392 00:25:42,290 --> 00:25:46,130 work, that's the most important part. How I'm feeling 393 00:25:46,450 --> 00:25:50,290 is going to unfold is like how 100, almost 100 years 394 00:25:50,610 --> 00:25:54,210 ago, Industrial Revolution unfolded, right? There were, for example, 395 00:25:54,450 --> 00:25:58,230 there were people who are making clothes by hand, right? 396 00:25:58,470 --> 00:26:02,310 Then the machines came and we got operators who would operate those machines 397 00:26:02,550 --> 00:26:06,310 to produce clothes at a humongous scale and see how 398 00:26:06,390 --> 00:26:09,990 many clothes we have today, right? Same thing is going to happen 399 00:26:10,550 --> 00:26:13,750 with AI in general. It's going to give us those machines 400 00:26:14,470 --> 00:26:18,150 which will increase the overall throughput of all the people 401 00:26:18,390 --> 00:26:22,070 out there for different professions. It's going to bring in a lot of prosperity. 402 00:26:22,940 --> 00:26:26,700 But at that time, people who were making clothes by hand, they went out of 403 00:26:26,860 --> 00:26:30,700 job. Maybe some of them learned how to operate the machines, but that's 404 00:26:30,700 --> 00:26:34,380 the important part. Learn to operate the machines. Like, don't just say that, hey, 405 00:26:34,620 --> 00:26:38,460 AI is going to take my jobs. Learn to use AI, embrace it. Your job 406 00:26:38,540 --> 00:26:42,220 is not going anywhere there. That's a really good way to put it. 407 00:26:42,620 --> 00:26:46,060 One of the examples I like is, um, there's self-driving cars, 408 00:26:47,180 --> 00:26:50,630 right? But there's also a lot of cars have adaptive cruise 409 00:26:50,950 --> 00:26:54,390 control. So for those not familiar with it, you know, it's basically cruise 410 00:26:54,630 --> 00:26:58,390 control, which will keep your speed, but it also has proximity sensors. So 411 00:26:58,390 --> 00:27:02,150 it'll apply the brake, right? And kind of keep you within a certain distance of 412 00:27:02,150 --> 00:27:05,430 the car ahead of you, et cetera, et cetera, et cetera. And also can keep 413 00:27:05,510 --> 00:27:09,270 you inside your lane. When I got that in, in 414 00:27:09,430 --> 00:27:13,110 my car for the first time, I felt like— It's a good analogy for AI 415 00:27:13,350 --> 00:27:17,060 because I'm still controlling the car. Right. I'm not fancy enough 416 00:27:17,140 --> 00:27:20,820 to have a full-on, you know, self-driving car, but, um, but I 417 00:27:20,980 --> 00:27:24,340 did notice that I became more like a captain of a ship, so to speak, 418 00:27:24,500 --> 00:27:28,180 where I would— I felt like I was guiding the car which direction to 419 00:27:28,260 --> 00:27:31,940 go. Now, in that case, I'm not a big fan of driving around 420 00:27:32,100 --> 00:27:34,580 all the time, but so it would be nice to have a full driving car. 421 00:27:34,900 --> 00:27:38,180 But I think that's kind of a good analogy for how, 422 00:27:38,980 --> 00:27:42,450 how jobs are going to look like, right? And if you're a software engineer, you're 423 00:27:42,450 --> 00:27:45,330 not going to be writing every line of code by hand anymore. I think, 424 00:27:46,610 --> 00:27:50,210 to your example, I think, you know, when humans were sewing, doing the 425 00:27:50,290 --> 00:27:53,810 sewing, right, individual stitch by stitch, clothes were made. And accordingly, 426 00:27:55,250 --> 00:27:59,090 clothes were expensive. And I think we're going to see that kind of happen with 427 00:27:59,730 --> 00:28:03,090 AI, right? Whether it's code, whether it's data engineering tasks. 428 00:28:03,810 --> 00:28:07,410 Because certainly I think anyone out there who is in a professional 429 00:28:07,970 --> 00:28:11,680 position where they're doing data engineering, There's plenty of 430 00:28:11,760 --> 00:28:14,560 things that are just kind of on the back burner that they would do to 431 00:28:14,560 --> 00:28:17,760 make their jobs more efficient, right? And I just think that, 432 00:28:18,640 --> 00:28:20,560 I think by having AI enabling, 433 00:28:22,480 --> 00:28:26,240 you have a lot more cognitive free room. And all those things 434 00:28:26,400 --> 00:28:30,000 are on the back burner, on the whiteboard somewhere, or kind of on a Post-it 435 00:28:30,080 --> 00:28:33,600 note somewhere about, hey, you know, I bet I can improve this process by 436 00:28:34,160 --> 00:28:37,780 X number of percent, but they're too busy keeping the machines 437 00:28:38,260 --> 00:28:42,020 working, right, keeping everything working. So with AI, I think you can have a 438 00:28:42,100 --> 00:28:45,940 bit of cognitive surplus where you can tackle those. Sorry, I cut you off. You 439 00:28:45,940 --> 00:28:49,620 were about to say something. No, nothing to add there. I think 440 00:28:49,620 --> 00:28:52,340 that's so spot on. This is where the world is going. 441 00:28:55,860 --> 00:28:59,620 That's cool. That's cool. So what's next for your company? Like, what, 442 00:28:59,860 --> 00:29:03,620 what's, like, what other big challenges do you think you 443 00:29:03,700 --> 00:29:07,340 all are going to take on? So one thing is, of course, 444 00:29:07,660 --> 00:29:11,500 right now we have been automating few of the data engineering use cases, 445 00:29:12,460 --> 00:29:16,220 but we are planning to expand to more and more use cases and at 446 00:29:16,300 --> 00:29:19,820 the same time improve the support of the data stacks 447 00:29:20,060 --> 00:29:23,500 we support as well. For example, you're talking about SQL Server, 448 00:29:23,900 --> 00:29:27,340 maybe older Hadoop, Hive stacks. There's so much information 449 00:29:27,740 --> 00:29:31,500 out there, and being early-stage company, we support certain 450 00:29:31,820 --> 00:29:35,080 stacks, certain stacks we don't support. So the plan is to expand 451 00:29:35,640 --> 00:29:39,400 into more ecosystems and at the same time expand for more 452 00:29:39,480 --> 00:29:43,320 use cases as well. Yeah, that's cool. Yeah, and it's funny 453 00:29:43,480 --> 00:29:47,320 because I know there's a lot of Hadoop legacy solutions out 454 00:29:47,400 --> 00:29:50,920 there. Somebody told me, and I don't want to pick on Hadoop, right? But 455 00:29:51,080 --> 00:29:54,760 Hadoop was one of the first, you know, big data solutions out there. And 456 00:29:55,400 --> 00:29:58,760 I would imagine there's a lot of— somebody told me that it was moved to 457 00:29:58,840 --> 00:30:02,400 the— what's Apache call it? The attic? the attic or the basement where it's basically 458 00:30:02,960 --> 00:30:06,320 considered a legacy product now. And, you know, obviously 459 00:30:06,880 --> 00:30:10,640 there's a lot of engineering— data engineering projects historically 460 00:30:11,120 --> 00:30:14,800 have not turned over that quickly, right? Like, you know, there were— there are probably 461 00:30:15,120 --> 00:30:18,240 batch jobs I wrote in the early 2000s that are still running, right? 462 00:30:19,760 --> 00:30:22,800 And that's not that unusual. So I think that 463 00:30:23,680 --> 00:30:26,970 because data modernization projects even if they're 464 00:30:26,970 --> 00:30:30,810 needed, they may not happen because people are too busy keeping 465 00:30:30,970 --> 00:30:34,570 the machines running, right? So if you can kind of— I think people are missing 466 00:30:34,730 --> 00:30:38,170 the point about AI taking their jobs. I know we're back on this again. You 467 00:30:38,250 --> 00:30:41,450 can always tell when the coffee hits me, right? Like, my mind goes in like 468 00:30:41,530 --> 00:30:45,370 3 different places at once. But the release 469 00:30:45,610 --> 00:30:49,290 of cognitive service, cognitive surplus, 470 00:30:49,450 --> 00:30:53,060 I think is going to do a lot of good things for The IT 471 00:30:53,460 --> 00:30:57,220 industry, and ultimately, I think, for the broader economy as a whole, right? 472 00:30:57,300 --> 00:31:00,980 I think you, you mentioned prosperity before, and I know it's a, a lot of 473 00:31:01,060 --> 00:31:04,820 people are very anxious about what AI can do. And you live, you know, 474 00:31:04,820 --> 00:31:08,100 you're in the Bay Area, you're probably at the epicenter of a lot of this 475 00:31:08,260 --> 00:31:11,220 back and forth, right? But I think that 476 00:31:12,100 --> 00:31:15,540 by releasing some of this, the human minds 477 00:31:15,940 --> 00:31:19,620 from the drudgery of some of the, the machinery that we have today, 478 00:31:20,350 --> 00:31:23,470 I think has enormous potential for improving IT, right? 479 00:31:24,430 --> 00:31:28,030 You know, take any kind of like regular person out there who, who has to 480 00:31:28,270 --> 00:31:32,110 interact with their IT department. Would they describe it as a wonderful experience, right? 481 00:31:32,190 --> 00:31:35,070 Would they describe it as a, you know, or is it more like going to 482 00:31:35,150 --> 00:31:38,830 the DMV, right? Probably more like the DMV. Although 483 00:31:39,070 --> 00:31:42,910 I will say credit where credit is due. In Maryland, I had to spend— I 484 00:31:43,070 --> 00:31:46,820 recently got a new car and it wasn't as bad as I feared 485 00:31:46,980 --> 00:31:50,020 it to be. You know, maybe they're using AI, I don't 486 00:31:50,020 --> 00:31:53,860 know. Yeah, yeah. And as you mentioned, right, a lot 487 00:31:53,940 --> 00:31:57,700 of this work we sort of parked it for later. Like you were talking 488 00:31:57,860 --> 00:32:01,380 about batch jobs or Hadoop jobs, etc. There's still a bunch of those things running 489 00:32:02,180 --> 00:32:06,020 because migrations had been nightmares, right? 490 00:32:06,180 --> 00:32:09,860 Like people spend millions for years. But now I think, now 491 00:32:09,940 --> 00:32:13,680 if I look at it with the angle of AI, AI can automate 492 00:32:13,840 --> 00:32:17,680 90-95% of those migrations as well. Now, who knows, like, even 493 00:32:18,000 --> 00:32:21,600 that whole modernization and digitization 494 00:32:22,080 --> 00:32:25,920 in the data world, that will start accelerating much faster because of 495 00:32:26,000 --> 00:32:29,840 this. And a lot faster, different companies can be on the 496 00:32:30,000 --> 00:32:33,760 most modern technology. And they can, they can take advantage of 497 00:32:35,120 --> 00:32:38,880 the newer features, the optimization technology. Yeah, I mean, 498 00:32:39,280 --> 00:32:42,720 because no, no company in their right mind is going to suddenly, you know, say 499 00:32:42,800 --> 00:32:46,060 we're going to hire, you know, we have this system, we can either 500 00:32:46,300 --> 00:32:49,900 A, keep throwing a little bit of money at 501 00:32:50,360 --> 00:32:53,660 it per year, right? And keep it going indefinitely. Or 502 00:32:53,980 --> 00:32:57,580 B, we're going to hire 100 new data engineers and all these project 503 00:32:57,820 --> 00:33:01,180 managers and things like that. We're going to spend billions on upgrading it. They're not 504 00:33:01,180 --> 00:33:05,020 going to do that, right? But they do need to upgrade, right? So then if 505 00:33:05,020 --> 00:33:08,780 it becomes a, well, you know, if you just spend just a little bit 506 00:33:08,860 --> 00:33:11,730 more than what you do have to do to keep the lights on, and you 507 00:33:11,810 --> 00:33:14,850 kind of have AI take a lot of the brunt work of that, right? Instead 508 00:33:14,850 --> 00:33:18,290 of hiring 400 people, you can probably hire, you know, say 509 00:33:19,010 --> 00:33:22,530 40, right, to do the migration. It starts to become more 510 00:33:22,610 --> 00:33:26,450 palatable, right? It starts to become more justifiable in a, in 511 00:33:26,450 --> 00:33:29,970 a balance sheet. And, and, you know, I think a lot of the breaches 512 00:33:30,050 --> 00:33:33,730 we've seen have been from improperly 513 00:33:33,970 --> 00:33:37,730 configured systems, right? Or legacy software 514 00:33:38,050 --> 00:33:41,250 that hasn't been patched or legacy software that has no business still running 515 00:33:42,130 --> 00:33:45,870 in this day and age. And I think you're right. I think AI is really— 516 00:33:46,110 --> 00:33:49,230 I think if you put the right type of glasses on, AI 517 00:33:51,710 --> 00:33:55,550 is a way to solve the problems that we've just 518 00:33:55,870 --> 00:33:59,550 been pushing back for years and years. And technical 519 00:33:59,950 --> 00:34:03,550 debt, that's the word I'm looking for. I think AI has the potential to 520 00:34:04,270 --> 00:34:07,950 be a great way to pay down technical debt. Yeah. 521 00:34:08,190 --> 00:34:12,039 But So while we're talking about 522 00:34:12,279 --> 00:34:15,879 that, like, what is the thing that your customers who are successful on your 523 00:34:16,039 --> 00:34:19,879 platform, what is the thing that they're most happy with at the end 524 00:34:19,879 --> 00:34:22,759 of the day? What are they like, they call you up and they say, wow, 525 00:34:22,839 --> 00:34:25,079 Pranesh, I'm so glad we did this because 526 00:34:26,679 --> 00:34:29,639 what is it, increased productivity? Is it something else? 527 00:34:31,079 --> 00:34:34,839 A few things, right? So for example, one example I gave was 528 00:34:35,159 --> 00:34:38,990 we can manage people's infrastructure using AI. Right? And 529 00:34:39,070 --> 00:34:42,350 we can optimize it. Now AI does it at scale 530 00:34:43,230 --> 00:34:47,070 at which humans can't even do it. For example, it can change the configurations 531 00:34:47,870 --> 00:34:51,630 300 times in a day. We can't do it, right? Like, how can we 532 00:34:51,710 --> 00:34:55,390 change the configuration of single machine 300 times a day? You need 150 533 00:34:55,550 --> 00:34:59,230 people doing just that. Yeah. And 534 00:34:59,710 --> 00:35:03,390 especially in the space of data, your workloads are always changing, right? 535 00:35:03,550 --> 00:35:07,300 As your data changes, your workload changes. So one of the things we have done 536 00:35:07,940 --> 00:35:11,460 is that infrastructure optimization, where we have built agents which can 537 00:35:11,780 --> 00:35:15,460 automatically analyze your infrastructure, change the configurations. We have built agents 538 00:35:15,860 --> 00:35:19,700 which can analyze the data pipelines and automatically optimize them. And 539 00:35:19,940 --> 00:35:23,300 of course, that saves tons of time for engineers, but at the same 540 00:35:23,460 --> 00:35:26,660 time, your infrastructure runs at 90 to 100% 541 00:35:27,300 --> 00:35:30,740 utilization, and those pipelines are always top-notch when it comes to 542 00:35:31,140 --> 00:35:34,740 optimization because people run millions of SQL queries and hundreds of thousands of 543 00:35:34,820 --> 00:35:38,660 pipelines. Now AI does that. The result is not just the 544 00:35:38,900 --> 00:35:42,740 engineering hours saving, but real dollar savings on 545 00:35:42,820 --> 00:35:46,580 the infrastructure cost also. So some of our biggest customers, we 546 00:35:46,660 --> 00:35:50,180 have saved them millions of dollars in their Snowflake and 547 00:35:50,260 --> 00:35:52,420 Databricks environment by optimizing this. 548 00:35:54,500 --> 00:35:58,260 And yeah, that's super happy about it. They're like, hey, tool pays for 549 00:35:58,500 --> 00:36:02,260 itself multiple times over. We are getting all these engineering productivity also. We're saving 550 00:36:02,420 --> 00:36:05,900 tons of engineering hours. But hey, right there and then you're saving me 551 00:36:06,220 --> 00:36:09,980 infrastructure costs also by automating that process by a huge margin. 552 00:36:11,660 --> 00:36:13,820 No, it's a great way to— that's a great way to look at it. And 553 00:36:14,140 --> 00:36:17,980 I really think there's so much opportunity. I know there's a lot of people, there 554 00:36:17,980 --> 00:36:21,660 are a lot of naysayers now about what AI looks like and, 555 00:36:21,820 --> 00:36:25,260 and, you know, will we, will we realize the gains that were been promised? 556 00:36:25,660 --> 00:36:29,340 I think we will. I think it's going to be in places where we may 557 00:36:29,420 --> 00:36:32,970 not think of, right? No one You know, and data engineers will be the first 558 00:36:33,050 --> 00:36:36,650 people to tell you, like, they're not usually first of mind for a lot of 559 00:36:37,370 --> 00:36:41,050 people in the C-suite, right? Data is like air. You don't think 560 00:36:41,290 --> 00:36:45,130 about it. Good data engineering is like air. You don't think about it 561 00:36:45,210 --> 00:36:48,970 until you don't have any, right? Like, it, you know, if you, you know, it's 562 00:36:48,970 --> 00:36:52,410 kind of lost in, in, in the haze, so to speak. And 563 00:36:52,810 --> 00:36:55,770 I think that, I think if people, 564 00:36:56,890 --> 00:36:59,770 organizations that get their data estates kind of sorted out, 565 00:37:00,700 --> 00:37:04,300 are at a competitive advantage to anyone that doesn't, right? And 566 00:37:04,620 --> 00:37:08,300 it's very often— that's why I always make a big deal in the intro is 567 00:37:08,540 --> 00:37:12,140 that people don't think about data engineering. I've been in hundreds of meetings where 568 00:37:12,940 --> 00:37:16,780 the data engineering aspect is often neglected, right? One 569 00:37:17,740 --> 00:37:21,420 story in particular was this guy had this great idea for this, 570 00:37:21,900 --> 00:37:24,700 that for his organization, and he was going to hire 571 00:37:25,910 --> 00:37:28,710 He was going to do all this sort of stuff, integrating all these different data 572 00:37:29,030 --> 00:37:32,870 sources, you know, from public, private, enterprise, you name it, 573 00:37:32,950 --> 00:37:36,630 right? You name it, he mentioned it, right? It was that type of guy. And 574 00:37:36,630 --> 00:37:39,750 he goes, I'm going to need 20 data scientists. And I'm looking at it, I'm 575 00:37:39,750 --> 00:37:43,350 like, you're going to need 18 576 00:37:44,550 --> 00:37:48,230 data engineers and 2 data scientists, right? Like, 577 00:37:48,550 --> 00:37:51,720 you know, like, you know, because like the data engineers, Quote unquote, 578 00:37:52,440 --> 00:37:56,280 the old school kind of proper data engineers, I 579 00:37:56,280 --> 00:37:59,640 mean, data scientists, they don't want to deal with SQL, right? They just want to 580 00:37:59,640 --> 00:38:03,400 get kind of their data. But like, in terms of the sheer 581 00:38:03,560 --> 00:38:07,000 amount of systems this guy wanted to integrate and do all the 582 00:38:07,320 --> 00:38:10,600 translation, I mean, he's going to need— and he's probably going to need, 583 00:38:11,480 --> 00:38:15,320 you know, 50 data engineers realistically, right? Or 20 of them doing 584 00:38:15,480 --> 00:38:18,690 the work of 50. But yeah, it just goes to show you 585 00:38:19,810 --> 00:38:23,570 that no one appreciates data engineering, right? 586 00:38:23,650 --> 00:38:26,690 You know, it's almost invisible. It's certainly invisible when it works 587 00:38:27,170 --> 00:38:30,770 well, and it's certainly visible when it 588 00:38:30,850 --> 00:38:34,370 doesn't work. But it's also when it's working kind of mediocre, 589 00:38:34,450 --> 00:38:38,210 it's also kind of invisible too, right? Yeah, yeah. 590 00:38:38,770 --> 00:38:42,130 A lot of times it gets underestimated how much time 591 00:38:43,110 --> 00:38:46,710 the work is required to put, I think, right amount of data 592 00:38:47,110 --> 00:38:50,710 on the table, right? And that's why a lot of times people don't 593 00:38:51,030 --> 00:38:54,550 understand why there is a huge backlog. Hey, I'm asking for simple data, why you 594 00:38:54,550 --> 00:38:58,230 are telling me it's too much? So there are 50 people ahead of you and 595 00:38:58,310 --> 00:39:01,990 every project is going to take 4 days, right? Takes time to find that data 596 00:39:02,310 --> 00:39:06,150 curated and put it somewhere where it can be consumed by, say, 597 00:39:06,470 --> 00:39:08,790 analytics use cases or data science use cases. 598 00:39:10,580 --> 00:39:13,380 You know, one of the things that comes up a lot is the joke is 599 00:39:13,540 --> 00:39:16,660 that DBAs used to stand for don't bother 600 00:39:16,820 --> 00:39:20,420 asking, right? But it— but I think that 601 00:39:20,740 --> 00:39:24,340 we had an earlier guest a long time ago now basically talk about that 602 00:39:25,380 --> 00:39:28,820 one of the important shifts, I think AI really puts the acceleration 603 00:39:29,620 --> 00:39:32,500 on the shift of mentality, is that data 604 00:39:33,140 --> 00:39:36,750 DBAs, Data engineering managers have to think of 605 00:39:36,990 --> 00:39:40,710 themselves less as gatekeepers and more like shopkeepers, 606 00:39:40,990 --> 00:39:44,670 right? They're not the bouncers at the door anymore, right? They're the people 607 00:39:44,910 --> 00:39:47,710 behind the counter that say, how can I help you? Bonus points if they can 608 00:39:47,790 --> 00:39:50,670 set up a system that is— again, I'll go back to the DMV, right? If 609 00:39:50,670 --> 00:39:54,510 you go to the Maryland DMV, I'm sure it's even more so where you 610 00:39:54,590 --> 00:39:58,270 live, given it's the Valley, right? There's plenty of self-service kiosks 611 00:39:58,430 --> 00:40:02,170 where if you just need a new XYZ, You can just scan 612 00:40:02,410 --> 00:40:06,250 your driver's license, enter some information, of course, 613 00:40:06,250 --> 00:40:09,850 enter some payment information as well, and you, you know, they'll, 614 00:40:10,010 --> 00:40:13,850 they'll print up or it's almost fully self-service. And I 615 00:40:14,010 --> 00:40:17,690 think that the data engineering organization in the enterprise in the future 616 00:40:17,850 --> 00:40:21,610 is going to look a lot more like that than in the way it's looked 617 00:40:21,770 --> 00:40:25,530 like historically. Yeah, a lot of it is going to 618 00:40:25,530 --> 00:40:29,290 be self-service also, right? People just come in if it's simple tasks. 619 00:40:29,370 --> 00:40:33,170 Hey, as an organization, data organization, will 620 00:40:33,330 --> 00:40:35,970 serve it for you. You don't need to bother us, or we don't need to 621 00:40:36,050 --> 00:40:39,650 bother you. Like, you can just get what you need. If it's complex stuff, 622 00:40:40,210 --> 00:40:44,050 then we will put our minions to work. Otherwise, our minions will work 623 00:40:44,290 --> 00:40:48,130 completely independently with you. Yeah. So do you 624 00:40:48,450 --> 00:40:51,810 have— do you track metrics in your product that says, 625 00:40:52,130 --> 00:40:55,810 hey, you know, this saves X number of hours, or is that something 626 00:40:56,050 --> 00:40:59,290 that Will come later? No, 627 00:40:59,290 --> 00:41:02,810 absolutely. We track it today. So it tracks what kind of tasks 628 00:41:03,370 --> 00:41:07,210 have been done already by AI for you, and we 629 00:41:07,370 --> 00:41:11,050 try to make sort of a guesstimate also. If this task was 630 00:41:11,210 --> 00:41:15,050 done by you manually without using AI, this is how much 631 00:41:16,570 --> 00:41:20,170 timewise it would have costed you. So that's one part of it. Second 632 00:41:20,490 --> 00:41:24,070 is we do a bunch of benchmarking. against some other 633 00:41:24,310 --> 00:41:28,070 tools out there. For example, if you use— if you try to do this 634 00:41:28,230 --> 00:41:31,830 task using just vanilla Claude Code or GitHub Copilot, or 635 00:41:32,310 --> 00:41:36,150 if you use some other AI solutions that Snowflake and Databricks has 636 00:41:36,310 --> 00:41:39,350 put out also, like Codex Code, or dbt has put out 637 00:41:39,510 --> 00:41:43,110 Visit, how much better we do against that also. So 638 00:41:43,430 --> 00:41:47,190 we do the benchmarking, we publish those benchmarks. That's great. There are open-source 639 00:41:47,350 --> 00:41:50,800 benchmarks around this. One is called Agentic Data Engineering, 640 00:41:51,360 --> 00:41:55,040 ADE, and the other one is Data Agent Benchmark, DAB. 641 00:41:55,280 --> 00:41:59,040 So we do that testing, we put out those benchmarks, and those keep continuously changing 642 00:41:59,360 --> 00:42:02,960 as the space itself is so fast evolving. Wow. 643 00:42:04,240 --> 00:42:06,640 It's not a real technology until there's benchmarks, right? 644 00:42:08,400 --> 00:42:11,920 Yes. In the data world, in infrastructure, we love our 645 00:42:12,080 --> 00:42:15,840 benchmarks, right? So why not do it? It's the easiest way to compare different 646 00:42:16,000 --> 00:42:19,850 tools and different options. Well, awesome. Uh, 647 00:42:19,930 --> 00:42:22,810 and the website is Ultima— Ultima— 648 00:42:24,010 --> 00:42:27,850 I'm sorry, an Ultima just drove by. I'm sorry. Let me 649 00:42:27,930 --> 00:42:30,170 help you there. It's called Ultimate, so 650 00:42:31,130 --> 00:42:34,890 ultimate.ai, and our open source 651 00:42:35,210 --> 00:42:37,770 project is called Ultimate Code, so ultimate 652 00:42:39,130 --> 00:42:42,890 and code. Just search on it and you'll find the repository. 653 00:42:42,980 --> 00:42:46,580 You should find our website also right there. Excellent. Thank you very 654 00:42:46,660 --> 00:42:50,500 much. And I appreciate— I know we had some issues rescheduling, so 655 00:42:50,500 --> 00:42:54,340 I appreciate your patience with that. And I'm really looking forward to 656 00:42:54,420 --> 00:42:57,700 checking this out. You have my curiosity because there's all these like little ideas that 657 00:42:58,100 --> 00:43:01,460 I have in terms of, you know, I have datasets and things like that, personal 658 00:43:01,860 --> 00:43:05,620 and otherwise, that yes, I would love to go through and 659 00:43:06,020 --> 00:43:09,130 organize those. However, I don't have time. Right? 660 00:43:10,010 --> 00:43:13,850 Just as a personal project. I'm telling you, when you start running your own home 661 00:43:14,010 --> 00:43:17,290 lab, you really start to feel the pain of what IT organizations 662 00:43:17,610 --> 00:43:21,450 feel, right? I have an LLM server, and when it's down, my 663 00:43:21,530 --> 00:43:24,650 kids tell me, and it just becomes like this, like, nightmare of 664 00:43:25,130 --> 00:43:28,730 situations. So, like, the best way to learn is to do. And 665 00:43:29,770 --> 00:43:33,530 I encourage everyone out there to build out a home lab with whatever you got 666 00:43:33,850 --> 00:43:37,130 and just start messing around with the AI tools to make things easier. 667 00:43:38,410 --> 00:43:41,650 And, uh, with that, I'll let the outro music play.