Speaker Application
SPEAKER APPLICATION
Sponsor
SPONSOR
Past Events
PAST EVENTS
2026 SF
2025 OAK
2024 AUS
2023 AUS
2022 AUS
2020 SF
All talks
A
L
L
T
A
L
K
S
Keynote
From Silo to Scale: How Data Infrastructures Evolved to Bring More Data to More People
Full abstract coming soon.
Speakers
Bhaskar Ghosh
, Partner, 8VC
View transcript
00:00
Welcome, BG.
00:01
What I'm really interested in a way is you went to Yale.
00:04
That's computer science.
00:06
Is that when you got out, what were your expectations?
00:08
What were you all hoping to do?
00:10
I had no idea.
00:11
I was going to become a professor at the University of
00:13
Illinois at Urbana.
00:16
This was post-Reagan, so there were huge grants coming to
00:19
these things called the supercomputing centers that
00:22
the kids here will not know about.
00:24
So basically Stanford, Caltech, Rice, Yale,
00:27
Urbana, and Michigan.
00:29
And a lot of the parallel, all the stuff you guys are using
00:34
now at the GPU level, it's funny to say this, a lot of
00:38
that stuff came from the parallel sparse matrix stuff
00:41
we did back in the day, the parallel blast stuff, and is
00:44
now coming back 30 years later.
00:46
But anyway, I worked a lot on parallel algorithms for
00:50
scientific problems and then accidentally got my PhD
00:53
around graph partitioning.
00:54
I was going to go teach at Urbana, and my advisor said,
01:01
you're not suited for academia, you're
01:02
much too impatient.
01:03
So go west, and don't work on hardware, work on software.
01:07
That's it.
01:08
And I had no clue how to apply it.
01:10
Accidentally, I had a paper at supercomputing, and there was
01:13
some company called Informix.
01:15
I had no idea what that was.
01:16
I'd never taken a database class.
01:18
So I would say that was the lucky break.
01:20
I ended up going to, at that time, the hottest
01:22
database company.
01:23
Remember, it was killing Oracle and Cyberset, and
01:25
Ingress at that time.
01:27
So for those of you who don't know, Informix had two
01:30
beautiful products.
01:31
One was an OLTP product, one was a
01:32
data warehouse product.
01:34
Amazing.
01:35
And then things went wrong on the CEO side, so the company
01:37
did not do well.
01:39
But yeah, I would say I was lucky to start my career in
01:42
data kernels.
01:43
I spent many years at Oracle in the RDBMS team.
01:47
And if you think of the team we had, imagine Snowflake came
01:50
out of our team.
01:51
Benoit Thierry and I wrote a lot of code together.
01:54
She started Qubole.
01:57
Vipul started Rubrik.
01:58
Dheeraj started Nutanix.
02:00
So those two floors have probably the most number of
02:04
unicorns coming out in infrastructure.
02:07
So I would say the schooling at Oracle was really important.
02:09
I went to Yahoo late.
02:11
I was very lucky to run this really hot ad
02:14
exchange called ByteMedia.
02:16
And we got to do some fairly deep distributed systems,
02:19
indexing, information retrieval work there.
02:21
We got to publish papers.
02:24
And then LinkedIn, five of us went to LinkedIn end of 2009
02:29
when LinkedIn was growing really quickly and didn't have
02:32
a very clearly defined data infrastructure strategy.
02:35
But we can dive into that more.
02:36
But I would say those two, I mean, Oracle's schooling in
02:40
RDBMS is what I'm most grateful for.
02:42
It's the most beautifully designed product, just took
02:44
the test of time.
02:46
And the engineering culture I saw at Oracle, the
02:49
meritocracy, was quality of code.
02:52
Everything was so amazing in the kernel team.
02:57
I'm extremely indebted to that phase of my life.
02:59
So I have a question for all you kids out there who may be
03:02
new to the data space.
03:04
How much do you think your kernel work has influenced how
03:08
you process what's going on and more the user side, the data
03:11
engineering and data analyst side?
03:14
Yeah, I mean, when you work, I would say if you start your
03:16
career in old time, like networking stack, if you
03:19
worked at AT&T, Bell Labs, or you worked at IBM,
03:23
DB2, or Oracle, these are very complex systems, some of the
03:27
most beautiful code ever written.
03:28
But the main thing you learn by being in a group like this
03:31
is separation of concerns.
03:33
You start thinking about architectural
03:34
stuff a lot more.
03:35
You start thinking about APIs a lot more.
03:37
We didn't call them APIs at the time.
03:40
We would call them functions and parameters, right?
03:44
I would say separation of concerns and systems thinking
03:48
actually is really important for both data infrastructure
03:52
building as well as data infrastructure businesses.
03:56
I think the thing you and I discussed about product
04:00
geometry, persona user geometry, budget geometry,
04:04
this thing of geometry or what you call lattice, I would say
04:08
a lot of those mental models came because we got to work on
04:12
such complex, multi-tiered, beautifully architected
04:15
software products.
04:17
Right.
04:18
Just in terms of those two framing, a way to think about
04:22
data, I just want to talk about that for a minute.
04:26
We're all familiar with the word stack.
04:28
And I think stack really became popular when the lamp
04:31
stack was something.
04:32
Because it really was a stack.
04:34
Things were piled on top of each other.
04:36
Which came out of Yahoo, I just want to call that out, for
04:39
the young people who have been there.
04:40
That's right.
04:40
And it's also funny how radical having an open source
04:45
stack was at the time.
04:47
It's hard to imagine it now, with open source so embedded
04:51
in the way we think.
04:51
But back then, it was like Martians had come and taken
04:55
over software.
04:57
But I think in the data space, what's changed is that, from
05:00
my perspective, this lattice metaphor works, where you've
05:03
got all these nodes, and then the edges are where processes
05:06
take place.
05:08
And then, BG has described these geometries.
05:11
And I'll let you explain it in a second.
05:13
And I think framing is really useful.
05:16
But the thing I want to say about that is, it's a layer of
05:19
abstraction using these metaphors.
05:21
And I think that that's been one of the things, as we
05:23
talked beforehand, that we'll probably reference more than a
05:26
few times, is that the increasing levels of
05:29
abstraction means that more and more
05:32
people can use the data.
05:33
So why don't you explain your geometries?
05:36
Oh, boy.
05:37
Do you want me to give a little bit of the LinkedIn
05:40
background, or do you want me to?
05:42
Why don't we do the geometries first, and then go into it?
05:44
Yeah, so jumping right ahead, the thing we think about as
05:48
builder, I mean, I also start companies as part of ABC, just
05:51
started a stealth analytics company, which uses Gen AI.
05:54
Something that we force ourselves at ABC to think
05:57
about is, first and foremost, product geometry and personal
06:00
geometry.
06:01
So product geometry is, suppose you want to build a
06:04
new product in the entire data lake, data warehouse,
06:07
analytics, ETL space.
06:09
It's good to go to a whiteboard and draw out
06:12
building blocks of what exists today.
06:16
Is there an ingest piece?
06:18
Is there an ETL or ELT piece?
06:21
Is there a substrata of a data lake piece or a
06:23
blob store piece?
06:24
Is there a BI tier?
06:27
Is there a SQL democracy tier like DBT?
06:30
Is there a warehouse sitting on the side?
06:32
Is there such a thing as reverse ETL?
06:35
When you draw all that out, you are forced to think about
06:38
how do they fit together?
06:41
How do they use each other?
06:42
How do they integrate with each other?
06:44
And then you'll start finding out that some pieces that
06:47
you're trying to build does not fit into it.
06:50
You have to start thinking about why doesn't it fit in.
06:52
Is it an API problem, usage problem?
06:54
Is it that you are putting too little functionality in?
06:59
Let me give you an example.
07:01
We started this metadata project at LinkedIn.
07:03
It was called something else in 2012.
07:06
And then Shashanka took it over and turned it into this
07:08
massive open source project called Data Hub.
07:11
And then from ABC, we founded this company called Acro.
07:14
So initially, you think of Data Hub as a data catalog.
07:17
Well, everybody builds data catalogs.
07:19
Not very interesting.
07:21
Then you start thinking, what's in that data catalog?
07:23
It's actually a metadata plane, which will have
07:25
information about tables.
07:26
It will also have information about pipelines.
07:29
Some data will have information about AI models
07:31
and their lineage.
07:33
Then you start thinking, is that product itself
07:37
a geometrical work?
07:38
The answer is no.
07:39
Nobody will pay for that metadata plane only.
07:42
They will start paying for workflows on top.
07:46
Then you need a search and discovery vertical workflow.
07:49
You need a data quality and observability workflow.
07:52
You need a governance workflow.
07:54
Maybe you need data lifecycle management.
07:56
So if you start thinking in that way, then what does your
07:59
MVP look like? How does it fit into it? Which personas use it? We have found that at 8VC and in our team
08:06
which does fairly heavy lifting in data infra, we found that a very useful mental model.
08:11
Now what we do is, you know, because a lot of my old friends now run really large data stuff at LinkedIn,
08:16
Pinterest, Airbnb, you know, all of us old-timers have friends everywhere.
08:19
I often go and kind of run this past them,
08:22
saying, you know, what does the product generally look like for you internally, even if you're not open sourcing it?
08:27
So it's been a fairly useful mental model.
08:31
Okay, great. Now you said that when we talked, you had said there were three. There's the product, product geometry,
08:38
budget, so first is persona geometry, second is product geometry, third is budget geometry.
08:45
I would say the first two things that you have to think about for go-to-market as a builder is
08:50
product geometry and persona geometry. If you don't get that right, it's very hard. Of course, budget has to come after that, but yeah.
08:56
Yeah, what's interesting too about the persona geometry from my perspective,
09:01
when I first started building data teams in the mid-90s and and beyond that,
09:06
we didn't separate out analytics from data engineering. Everyone did everything,
09:11
but over time, particularly as data got
09:14
larger, that you needed a little more specialization in it.
09:19
So having established this way of thinking about things, let's talk about your experience at LinkedIn.
09:26
What a cauldron of
09:29
projects. We were just trying to survive day-to-day, so we got there. Most of us, you know, I was very lucky. In my entire career,
09:35
I'm always lucky to work with people smarter than me, and I always worked on data. Those are the two formulas
09:40
I've seen. Hard formulas, but if that's true, then you'll probably learn a lot. When we got there,
09:46
we were like at 50 million unique users in the graph, which at that time we thought was huge. When I left
09:52
fall of 2014, it was
09:55
575 million. So, you know, went through that level. Facebook saw even more growth, but LinkedIn as a source of truth graph was hard.
10:02
So the first thing we had to work on was the online side, which this conference is a bit more analytics and data lake focused.
10:09
The source of two databases, I don't say a lot, but anyway, so I'll keep that part short.
10:14
So much of LinkedIn at that time was running on Oracle, and
10:19
Sean Luke, who was the founding CTO of LinkedIn, you guys who are interested in history, wrote an amazing paper at QCon
10:26
2007 or 8,
10:28
about change area capture by reverse engineering Oracle redo and undo logs. Amazing. That guy was a genius.
10:34
So we had this thing that was called Data Bus, and we had a bunch of Oracle instances, and everything was falling over.
10:41
And I
10:43
was lucky to hire a lot of smart people. Shashanka Swaroop, Kishore Sweekem, Jay Krebs was already there
10:50
before me.
10:51
We had no idea what to do.
10:53
So we kind of started looking at MongoDB very deeply. We brought in MongoDB, and it cleared over within two days,
11:00
which is okay. Mongo is a beautifully designed product,
11:02
way ahead of its time with the whole document style
11:06
DSL, which is really important.
11:09
Kind of did something that old-timers say Yahoo should have done with the LAMP stack, but Mongo did it first.
11:15
So we looked at it, and then we were lucky to go to either VLDB or SIGMOD.
11:21
Because you mentioned Jeffrey Dean, I wanted to tell you this. We were at Arizona,
11:26
no, we were in Indianapolis at a SIGMOD or VLDB, and you know, this is like first three months at LinkedIn.
11:31
I have no clue. I've hired these really smart people. We have no clue what to do.
11:36
We know that we have to build something with the document style programming and data model, API.
11:42
So I talked to Jeff Dean, basically like going to
11:46
Moses and Muhammad, going to the mountains saying, what should we do? He gave a really good piece of advice.
11:51
He said, do not build the backend from scratch.
11:54
Use something off the shelf, which is open source. That was like amazing advice.
11:59
So we ended up building what turned out to be the biggest piece of data infrastructure at LinkedIn that nobody has ever
12:05
heard about. It's called Espresso.
12:07
It's a source-of-truth database.
12:10
Horizontally scalable. We built it initially with the MySQL InnoDB backend for the transaction part, and a MongoDB-like
12:18
programming API on top. And you guys will be surprised to hear, 80 to 90 percent of all of LinkedIn still runs on,
12:26
I would say, thousands of instances of Espresso, supporting 5 plus million QPS.
12:31
I would say that was the heaviest lift.
12:33
I would say the Espresso journey was the hardest journey for LinkedIn data infrastructure on the online side.
12:38
And then Jay, Neha, Jun, and team did an amazing job with Kafka.
12:43
Data bus, which is the change data capture thing, you know, we just wrote a check into a
12:49
Y Combinator company called PRDB. Change data capture is still a very hard problem. Bringing change data logs
12:57
to where it needs to be used is very hard. So that's why we built data bus. I remember Shashankar wrote a beautiful paper back in
13:04
the day, but we merged it with Kafka finally.
13:07
So and then, you know, we of course had a cache, and then we built a very good graph engine, which kept on getting updated.
13:14
I didn't build it. Another team did. But I would say those were the four key pieces of the online side.
13:19
So that was the online journey. We just want to keep it brief, but I'll hand the mic over to you.
13:23
You know, what other you want to... Sir, one of the things I want to talk about is this fits in with this pattern.
13:29
I think you'll hear a few times.
13:31
Why were they building this?
13:33
No one was providing these things.
13:36
And
13:37
open source has its issues, and we can talk a little about that. It is good at functionality.
13:43
It's kind of pared down, usually, because someone is doing a point project.
13:49
I think, and I don't know if Julian Ledem is here,
13:51
but I think he's an exemplary open source project person, and he did Apache
13:57
Arrow and Parquet, and I think he really did a great job of reaching it horizontally.
14:03
But open source leads to a lot of fragmentation,
14:06
because now you've got a lot of piece parts, because all these people are solving these problems.
14:11
But they do sometimes coordinate, and he used a great term, and it was the spirit of
14:17
service. Is that what you thought?
14:19
Spirit of service.
14:20
That why would you build this when you're at LinkedIn, when you're not going to be a company selling these products?
14:27
Well, it's because you don't, you can't buy it anywhere. Why open source it?
14:32
So instead of just your team, you can have tens, hundreds, thousands of committers, perhaps,
14:38
helping you refine your product and taking it to the next level. It's a great virtuous circle in a lot of ways, and
14:46
I think that's why open source has become so important to the world,
14:51
to it. But you seem to kind of live this spirit of service.
14:55
Maybe you can explain how that affects...
14:58
Open source, I want to call something, Jonathan Goldman, you mentioned him. That guy was amazing.
15:02
He's the guy who came up with this
15:05
amazing thing called People You May Know.
15:07
You know how you densify the graph? That Facebook.
15:10
We should have patented and got an IP on it. He did it before we got to LinkedIn, so amazing guy.
15:17
Spirit of service here, I think.
15:20
It's funny, as a VC, finally you find out that you don't have control of anything. Nobody will listen to you.
15:26
So you're ultimately here to serve the entrepreneurs and maybe nudge them a little bit, right?
15:31
If you build data infrastructure or AI infrastructure, you build horizontal disciplines in an enterprise,
15:37
ultimately, nobody would want to listen to you because they're busy shipping their product.
15:42
And I think the personality of evangelization that you have to take on is a very, very light touch one.
15:48
Steeped in a spirit of service because you're trying to make the product side successful.
15:53
In very infrastructure heavy
15:56
culture companies like Google, maybe you don't need to. You just can go. But I would...
15:59
Say, something we did very well at LinkedIn was actually
16:05
evangelize stuff, get people to adopt in the right way
16:08
slowly, and not railroad them.
16:11
And we talked about spirit of service, and influence, and
16:15
evangelize rather than force as a very, very key part.
16:18
And when we were building infrastructure, I remember
16:22
when we built Espresso, the amount of time we spent with
16:24
the feeds team, the profile team, the mail team, which
16:28
were major teams at LinkedIn who came before us, seriously
16:31
good engineers.
16:33
We had to spend a lot of time with a lot of humility
16:36
listening to them.
16:37
And that was very well done.
16:41
And you know what ended up happening with the
16:43
programmability, the DSL side of the online side is because
16:46
we listened to them with a lot of humility, the data model
16:50
design of Espresso was amazing, the extensible
16:54
JSON-based data model design.
16:56
So humility might even lead to good design.
16:59
That could be one.
17:00
Right.
17:01
I actually think open source actually helps with the whole
17:03
humility equation.
17:05
Yes.
17:05
Open source, at that time, we came from Yahoo.
17:08
Yahoo was just a spectacular place.
17:10
Yahoo hackathons were legendary.
17:13
I mean, during my grad school days, I first learned about
17:15
open source from this crazy guy named Richard Stallman, who
17:18
would come to Yale to talk about it.
17:20
Too many young people here may not know about it.
17:23
Anybody who's hacked Emacs Lisp, which you and I did, we
17:27
owe a huge debt of gratitude to Stallman.
17:30
But then Yahoo was a cauldron of open source, just open,
17:32
very democratized stuff, sharing.
17:36
During my time, there were probably 50-plus
17:38
implementations of distributed hash tables inside Yahoo.
17:41
Can you imagine?
17:43
That people were using with each other, which is crazy.
17:47
DHTs are very hard to implement.
17:49
I would say the Hadoop thing was amazing at Yahoo.
17:53
Sadly, with two companies being formed, all of that stuff.
17:57
When I went to LinkedIn, LinkedIn was just beginning to
17:59
adopt Hadoop and became probably the biggest user
18:03
deployed enterprise use case of Hadoop, other than Yahoo.
18:07
It was massive.
18:07
And now it's still there, slowly getting out of Hadoop.
18:11
Open source at LinkedIn was part democracy,
18:13
part we had to compete.
18:16
We had to go hire engineers who had offers from Google,
18:20
Facebook, PayPal, eBay, all great companies at the time.
18:24
And we found that open source was also a hiring mechanism.
18:28
It was also a cultural mechanism.
18:30
And now I realize that open source is also a quality
18:33
mechanism.
18:34
If you're doing open source, you will get found out.
18:38
If you don't do a good job, the community doesn't have
18:41
anything to lose by telling you that you're wrong, right?
18:43
So I would say all of that, most of 80% of that, we got
18:47
right at LinkedIn.
18:48
And it continued right, left, end of 2014.
18:50
But I think it's continued.
18:51
They've done amazing open source stuff.
18:52
I think the whole Iceberg-based thing called Open
18:57
House, they've come up with this beautiful bunch of stuff.
19:00
A lot of good open source stuff.
19:01
I mean, I think LinkedIn even committed to pretty
19:05
sophisticated shuffle stuff for Spark.
19:08
I forgot what the project is called.
19:09
I was looking at it.
19:11
So open source was a cultural aspect, a hiring mechanism.
19:17
Now what's the flip side?
19:19
Flip side is once you are actually building stuff and
19:23
shipping to production, how do you maintain the balance of
19:28
outward-facing time versus production time, right?
19:32
Because open source also has a lot of seduction about it.
19:35
There's a community.
19:36
You go and give talks, right?
19:39
You enjoy the interaction with other people, with other
19:42
smart engineers.
19:43
And shipping stuff to production, fixing P0P workbox
19:47
is boring and hard.
19:49
So I think from a management and leadership perspective,
19:51
that balance has to be struck, right?
19:53
That's number one.
19:54
Number two is you mentioned fragmentation.
19:57
And I'll just make a brief comment.
19:59
I think data usage, data type, exploded in
20:03
most large enterprises.
20:05
Facebook was amazing.
20:06
I think Facebook just did phenomenal work in all parts
20:09
of data.
20:10
LinkedIn, surprisingly, we did.
20:11
And it's continued.
20:12
The tradition has continued for the last 10 years.
20:15
I left 10 years ago.
20:16
But what happens when you do open source and when you have
20:19
an enterprise where data is a first-class citizen, where
20:23
profit is tied to data, not just cost?
20:27
People tend to build a lot of bespoke stuff.
20:30
And now you go and open source it, right?
20:32
So what happens outside?
20:34
So that's a much harder question.
20:37
Would a company survive there?
20:40
Can a VC invest in it?
20:41
So I would say the same thing that a lot of us have done,
20:43
and the community I was part of has done, is we have
20:48
open sourced a lot of stuff, which is a good thing.
20:50
But the supermarket has also become very large.
20:53
So I think, for good or for bad, there's a lot
20:56
of fragmentation.
20:57
So now you start thinking about it from a user perspective.
21:00
If you're the head of data engineering or IT at a big bank,
21:05
how many things are you gonna bring in?
21:06
Are you gonna bring in 15 moving parts?
21:08
Or are you gonna pay for four, right?
21:10
And that's one part.
21:12
The other part that VCs are having to think about is,
21:16
open source projects give you one huge unfair advantage,
21:18
which is the bottom-up go-to-market network effects,
21:21
if it's done well.
21:22
If you look at Kafka, you look at Clickhouse,
21:24
look at DBT, you look at Data Hub.
21:28
Fair enough.
21:30
So the distribution advantage is there.
21:31
The question is, how do you make sure that you can rank
21:36
open source projects as an investor and builder
21:39
as something that would become economically viable?
21:41
I think that's a fairly hard problem.
21:43
I'll just leave it at that.
21:44
Yeah, I wanna, in a way, take off from that.
21:49
The tech hubs, Boston, San Francisco,
21:52
Seattle, New York, Boston, tend to attract
21:56
a very high level of dedicated engineers
21:58
who are interested in solving particular problems.
22:01
But there's a huge world beyond them.
22:04
And I think sometimes those folks
22:06
with a very tight engineering focus
22:09
don't pay enough attention to the users.
22:11
And they think all users are as sophisticated as they are.
22:15
And they aren't.
22:17
They have just a less engagement,
22:19
they didn't get the academic training or whatever.
22:21
I'm curious, now as you're an investor
22:23
and you're thinking about things around open source,
22:27
how much does thinking about the user
22:29
and the use cases matter in terms of your ability
22:32
to make decisions and figure out what to bring?
22:35
Can we talk about the DSL part of this a little bit?
22:38
I think Raja and I had an amazing couple of discussions.
22:41
We only know each other for two weeks now.
22:43
We've been finishing each other's sentences.
22:44
It's been amazing.
22:47
Two old fogeys, yes.
22:49
But we've been lucky to do what we have.
22:51
Timing for our careers has been amazing, no?
22:55
So user, I think the one thing Raja and I talked about
22:58
that's fascinating for all you guys
23:00
who want to build companies is,
23:02
what is the lingua franca, right?
23:04
I mean, from my first half of my career,
23:07
the lingua francas that ruled were SQL in the bottom half
23:10
and PL SQL, something procedural on top.
23:13
And then your SQL will let you do more complicated things,
23:19
manage sessions, open cursors, all of that stuff,
23:22
which a SQL query declaratively,
23:25
it let you manage state, right?
23:27
So that will never go away.
23:30
If you look at what's happening
23:31
in the Python, Jupyter Notebook, Pandas area,
23:33
what it lets you do is really to manage state
23:35
and be sticky, right?
23:37
That you do something, publish something,
23:39
do something, publish something, right?
23:41
So we started thinking, what about SQL?
23:43
Because we all thought SQL would die 15 years ago.
23:46
That's what it's worth.
23:47
I never did.
23:48
I got into it.
23:49
So what has happened?
23:50
Any of you Hadoop people here remember me arguing
23:52
about the lack of SQL?
23:54
So funnily, the Hadoop community,
23:58
especially Joy and Ashish from.
23:59
who invented Hive, huge debt of gratitude
24:03
because they ultimately said we can make SQL work
24:06
directly on top of MapReduce and Blobstores.
24:08
That's like unthinkable, right?
24:10
Remember at Yahoo, we went the other way,
24:11
we went the PL-SQL route and built Pig.
24:15
By the way, Pig and Hive both stood like
24:17
got the test of time award at SIGMON.
24:19
It's kind of important.
24:20
I hope the young practitioners here pay attention to it
24:23
because, you know, a lot of sawtooths happen, right?
24:26
SQL, I would say for Hive, and then I want to call out DBT.
24:30
DBT is amazing because it started really
24:33
codifying SQL as code.
24:35
You can check in SQL, you can share SQL,
24:37
you can almost share views in SQL, right?
24:41
So SQL actually exploded in the last 10 years,
24:43
whereas we had expected SQL to die, right?
24:46
So it's important from a user perspective
24:48
for anybody building or thinking of users
24:51
who think about what language will the program in.
24:53
I would say the second thing that happened is
24:55
at LinkedIn, when I was there,
24:58
all the fundamental ETL pipelines, okay, ELT,
25:01
we're not going to argue about that right now,
25:04
were built in Java.
25:05
Tower of Hadoop.
25:06
So today, I would say,
25:10
because of Iceberg and because of Arrow
25:14
and similar projects,
25:16
you have a polyglot set of stuff running on top, right?
25:19
You can have Python, you can even, God forbid,
25:21
you can write in C or Rust, you shouldn't.
25:24
Maybe even Java.
25:26
So there's a polyglot set of stuff.
25:28
Now, if you look at those languages
25:30
actually running in production,
25:32
I don't think the user community is that big.
25:34
It's a bunch of data engineers, right?
25:36
So it's important to think about that,
25:38
that if you're going to build a company below,
25:40
what is this community going to do?
25:42
So that's number two.
25:43
So one is SQL, one is this.
25:44
Then you start thinking about the revolution
25:46
that our old friend DJ,
25:50
who's coined the term data science,
25:51
along with Jeff, right?
25:53
I would say the data science community
25:55
has brought a lot to the table
25:57
in the experimentation phases, especially, right?
25:59
Again, borrowing from PL-SQL,
26:01
session-based work, Jupyter Notebook, and Python.
26:04
Python has become a lingua franca there.
26:06
So that is not going away.
26:08
So now you have that community.
26:10
So if you want to think about that user geometry,
26:13
what will that do?
26:15
We did do a company called Ponder,
26:17
which was the Python pandas,
26:19
with a project that's called Modin,
26:20
open source community that came out of Ice Lab,
26:22
which we sold to Snowflake last year.
26:27
If you guys want to do companies or build stuff
26:29
for that user community,
26:30
which is in the experimentation phase,
26:32
maybe even the serving phase of models,
26:35
think deeply.
26:37
Is that community going to be able to pay you
26:39
a lot of money or not?
26:40
I'd leave that as an open problem.
26:42
Sure, so time flies and you're having fun,
26:46
running out of time.
26:47
I do want to, I hope people have a couple questions.
26:50
Just wanted to say a couple key takeaways
26:53
that as we talked about this beforehand,
26:55
I just want to kind of go over.
26:56
Hopefully it'll be something that you get from what's here.
27:00
One is as an expression that history
27:02
doesn't repeat itself, it rhymes.
27:05
And I think from our perspective here,
27:08
being in the space a long time,
27:09
we're seeing a lot of rhyming.
27:11
There's things that are going on now
27:12
where there's precedent before.
27:14
It doesn't mean it's going to be exactly the same,
27:16
but you have a context for thinking about it.
27:18
So it's worth knowing history.
27:20
Jamak showed the original Unix paper.
27:23
If anyone told me afterwards about remember the 70s,
27:26
I've got a great thing about that.
27:28
Two is just being curious.
27:30
Luckily, the data space encourages that being curious.
27:34
Watch what folks are building when the thing's not available
27:38
because there's a market need.
27:39
I mean, it's so important to see that.
27:42
And you mentioned humility.
27:45
It is the best way to learn is to figure out you're wrong.
27:50
And focus, change, and learn from that.
27:54
♪♪