searchlore

Back to Resource

All Segments

a16z Podcast | From Data Warehouses to Data Lakes

a16z Podcast | From Data Warehouses to Data Lakes

24 segments available

From the silver age of on-prem software companies like SAP and Siebel Systems to the golden age of enterprise software-as-a-service, we're now seeing an explosion of data. All types, all sizes, and all over the place. And much of it is a sort of industrial "data exhaust", where companies aren't quite sure what question to ask of the data but are being bombarded with data due to the variety of data sources available today -- from websites to sensors (and therefore data capture) everywhere. Before there is even a signal in the noise. So how do you solve a problem like this-Data? Beyond requiring new types of plumbing and integrations, enterprises now expect -- given the age of mobile, web, cloud, and heck, let's add millennials to the mix too -- self service. To be able to ask, get, fit (curve-fit), predict. To take back the enterprise from the patchwork of integration and number of vendors we all have to deal with -- the scope of which most companies in fact are not truly aware of. It's about the lifecycle of data in the enterprise, argues Snaplogic founder and CEO Gaurav Dhillon in this episode of the a16z Podcast, in conversation with Scott Kupor. It's in fact about the evolution of data overall -- from data warehouses to "data lakes": in stages, from purification (like wrangling data) to bottling (prepping for consumption by data scientists) to making sense of streams and streams of data!

Segments Timeline

1
0:00 - 1:14
1:14 duration270 words

The Evolution of Enterprise Application Integration

Gaurav Dhillon discusses the early days of enterprise application integration (EAI) in the late 90s, highlighting the challenges faced by companies like SAP and Siebel Systems. He explains how businesses were struggling with business process reengineering and the need to integrate various applications, emphasizing the complexity of replacing core applications like finance.

"this is Scott Cooper I'm here with Gaurav Dillon the founder and CEO of snap logic and thank you Cora for joining us pleasure to be here Scott we're gonna cover a bunch of topics and keep this as a ve..."

2
1:14 - 2:36
1:22 duration310 words

The Role of Middleware in Enterprise Integration

Dhillon elaborates on the role of middleware in connecting applications within enterprises during the late 90s. He explains how integration was often subservient to major applications from vendors like Oracle, and how the limited audience of specialized users impacted the integration landscape before the web era.

"bring that stuff in do you mean data so it was there information sitting in those applications that people needed to kind of be able to talk to one another as part of this range here and what was the ..."

3
2:36 - 4:01
1:25 duration302 words

Architectural Shifts in Data Management

The conversation shifts to the architectural changes in data management from the late 90s to today. Dhillon highlights the transition from installing applications to using web-based solutions, emphasizing the profound impact of the internet on data types and the need for new integration strategies.

"limited audience of usage right remember this is all before the World Wide Web boy it was even before the web right this is the mid to late 90s so we things have changed quite a bit right to drill dow..."

4
4:01 - 5:30
1:29 duration281 words

Self-Service Expectations in Modern Enterprises

Dhillon discusses the rise of self-service as a critical requirement in modern enterprises, contrasting it with the past where specialized teams handled reporting. He notes how companies like Salesforce have pioneered self-service capabilities, fundamentally altering user expectations and the integration landscape.

"the web what does this do is you're no longer building installing applications you are essentially using websites you don't just use the Internet to buy your books you use the Internet to balance your..."

5
5:30 - 7:00
1:30 duration313 words

Proliferation of Applications in the Enterprise

The segment focuses on the explosion of applications within enterprises due to the rise of SaaS and web-based solutions. Dhillon explains how the traditional limitations on the number of applications have been lifted, leading to a significant increase in the number of tools used by business-level users, often without centralized IT oversight.

"one other thing that I also at least I've always thought is different is proliferation of applications in the enterprise so yes I agree we've got new data types we have new data sources right they're ..."

6
7:00 - 8:36
1:36 duration321 words

Rethinking Data Integration Strategies

Dhillon concludes by addressing the need to rethink data integration strategies in light of new data types and self-service models. He emphasizes the importance of adapting to modern data plumbing requirements and the shift from traditional integration methods to more agile, web-based approaches.

"shocking is people think it's do X right large enterprises they've layered on all these new applications many of which are now web-based applications at the same time probably we haven't retired many ..."

7
8:05 - 9:20
1:15 duration263 words

The Rise of Self-Service Data Integration

The conversation shifts to the importance of self-service in data integration. Dhillon highlights how traditional integration methods have been hidden away in enterprises, and the demand for self-service tools is becoming critical. He discusses the benefits of empowering business users to access and analyze data independently.

"you're no longer using rows and columns you're using a document model the way the worldwide web works the way browser talks to a website is the way these applications all function so I think that's a ..."

8
9:20 - 10:10
0:50 duration192 words

The Virtuous Cycle of Data Production

Dhillon describes the lifecycle of data in enterprises, emphasizing the continuous production, management, and consumption of data. He notes that as organizations recognize the benefits of analytics, they are motivated to capture more data, creating a self-reinforcing cycle of data utilization.

"to sell serve themselves okay so those two things are I think is sort of what's what's causing this to be a huge problem that makes a lot of sense so we've been talking about aggregation of data and t..."

9
10:10 - 11:12
1:01 duration217 words

Evolving Analytics: From Reporting to Predictive Insights

The discussion transitions to the evolution of analytics, moving from traditional reporting and business intelligence to predictive analytics. Dhillon highlights the shift in expectations from businesses, driven by technology companies that utilize advanced analytics and machine learning to forecast trends.

"healthy trend I want to shift gears a little bit we've been talking about you know movement of data right and integration of data across applications but I'd love to kind of shift the conversation a l..."

10
11:12 - 12:20
1:08 duration221 words

Data Lakes vs. Data Warehouses

Dhillon contrasts the traditional data warehouse model with the emerging concept of data lakes. He explains how data lakes facilitate modern analytics by allowing data scientists to access and analyze vast amounts of data, enabling more sophisticated predictive modeling and insights.

"those are discovery engines recommendation engines machine learning artificial intelligence algorithms although I think AI is often frequently misused I think missile machine learning a police thinks ..."

11
12:20 - 13:31
1:11 duration201 words

The Role of Data Scientists in Modern Analytics

In this segment, Dhillon elaborates on the role of data scientists in the analytics landscape. He discusses the collaboration between data scientists and business users to develop predictive models, emphasizing the need for data engineering to support these efforts in a rapidly evolving data environment.

"results that emerge and for that you continuously need to dip into the data lake bring forth data and you need to iterate on the algorithms so so this is what is happening and this is causing a profou..."

12
13:31 - 14:30
0:59 duration201 words

The Impact of Computational Advances on Data Science

Dhillon highlights how advancements in computational power and reduced costs of storage have transformed data science. He discusses how these changes have enabled more complex algorithms and predictive analytics, leading to significant shifts in how businesses leverage data for decision-making.

"people who are doing predictive analytics need to iterate the data with the algorithms so you divide up that task into a on the one hand a data scientist who is talking to the business person trying t..."

13
14:30 - 16:06
1:35 duration323 words

The Historical Context of Data Warehousing

The segment concludes with a historical overview of data warehousing, tracing its evolution from the 1990s to the present. Dhillon discusses the initial challenges that led to the development of data warehouses and how they have adapted to meet the needs of modern businesses, setting the stage for the transition to data lakes.

"use two words data warehouse the old you use this exciting new term data Lake that's right so I just let's kind of unpack those a little bit you know what was the problem again five ten years ago when..."

14
16:00 - 16:50
0:50 duration167 words

The Need for Data Lakes

As data needs evolve, Dhillon argues that traditional data warehouses are no longer sufficient, especially for the millennial generation and in marketing. He discusses the growing demand for AI and real-time data processing, which drives the shift towards data lakes.

"that was what that term is and look they're still running and lots and lots of places and you get a good historical perspective but that is no longer enough you know that is they particularly in this ..."

15
16:50 - 17:30
0:40 duration134 words

Understanding Data Lakes

Dhillon elaborates on the concept of data lakes, describing them as a repository for all types of data without duplication. He explains the importance of retaining data for future analysis, even if it appears to be noise at present.

"the packaged furniture industry for example initially you go and get planks of wood you saw them and you make you know furniture and then somebody says oh we'll have furniture pre-made and you can pai..."

16
17:30 - 18:10
0:40 duration134 words

The Evolution of Data Types

The discussion continues with Dhillon highlighting the diverse types of data that data lakes accommodate, including JSON, web exhaust, and machine data. He contrasts this with the simpler data landscape of the past, emphasizing the exponential growth of data sources.

"as computers get faster which we know happens every 18 months they double you're going to be able to get signal out of the noise right so so did the data leak fundamentally is about more data of more ..."

17
18:10 - 19:00
0:50 duration172 words

Data Wrangling Challenges

Dhillon addresses the challenges of wrangling vast amounts of data in a data lake environment. He outlines a three-stage model for data processing, from raw data to purification and bottling, which is essential for making data usable for data scientists.

"restocking because people said wait a minute that could be useful to me to know what's going on in what store how people are buying this and correlations and so on and so on so and now we have that on..."

18
19:00 - 19:40
0:40 duration149 words

Architectural Shifts in Data Analytics

The conversation shifts to the architectural changes in data analytics, highlighting the need for a hybrid approach that combines legacy systems with modern data lakes. Dhillon discusses the importance of real-time data processing and the role of data scientists in this new landscape.

"we're seeing a lot of companies adopt that because they need a organizing principle they need a place to hang their hat they need to think about how to deploy these technologies industrially how to go..."

19
19:40 - 20:30
0:50 duration200 words

The Future of Data Lakes

Dhillon speculates on the future of data lakes, suggesting they may evolve into cloud-based solutions. He discusses the potential for data lakes to integrate with cloud services like AWS and Azure, enhancing the speed and efficiency of data analytics.

"aggregation of and creation of the data lake itself and then on top of that you would put what reporting tools like a tableau or something else like that would be the mechanism to actually surface som..."

20
20:30 - 21:10
0:40 duration129 words

Real-Time Data Streaming

The discussion concludes with a focus on the shift from batch processing to real-time data streaming. Dhillon emphasizes the importance of immediate data access and the implications for various applications, including advertising and sales.

"how are people consuming it is going to be through certain Web Apps and and sometimes you're just putting that signal into an app to make a decision we're seeing precursors of this for example in AD t..."

21
20:51 - 22:14
1:22 duration284 words

The Evolution of Data Lakes

Gaurav Dhillon discusses the transition from data lakes to potential data oceans, emphasizing the need for architectural shifts in data management. He highlights how cloud technology and big data are converging, suggesting that data lakes may evolve into cloud formations, enhancing analytics capabilities and enabling real-time data processing.

"of thing is gonna happen for lots and lots of applications and you're starting to see companies like Salesforce for example make investments in machine learning to improve the efficacy of your sales p..."

22
22:14 - 23:39
1:25 duration303 words

Hybrid Data Environments

In this segment, Dhillon explains the hybrid nature of data environments, where legacy systems coexist with modern cloud applications. He illustrates the importance of integrating various data sources, including marketing data and machine data, into cloud-based data lakes to optimize data usage and analytics.

"cloud formations you actually might have data lakes exist in Microsoft Azure or in AWS as s3 buckets or we'll see what Google does with bigquery BigTable that that is going to give you in in a sense a..."

23
23:39 - 25:01
1:22 duration322 words

Self-Service Data Analytics

Dhillon emphasizes the shift towards self-service analytics in enterprises, comparing it to the evolution of music consumption. He argues that businesses should reclaim control over their data processes, enabling users to create customized data solutions without relying heavily on IT departments.

"put your marketing data into a data Lake in the cloud with Azure because you are going to have your website be the primary producer of data your on-premises days is actually very small and then you ha..."

24
25:01 - 27:53
2:51 duration608 words

Changing Roles in IT

This segment explores the evolving roles within IT organizations, particularly the collaboration between CIOs and CMOs. Dhillon discusses how technology has become integral to business strategy, leading to a more integrated approach to managing data, security, and cloud computing in large enterprises.

"bit for for folks we've kind of almost gone full circle and hopefully we're going even farther down the circle which is you know as you described we've gone from kind of a world of the application ven..."