Speaker Application
SPEAKER APPLICATION
Sponsor
SPONSOR
Past Events
PAST EVENTS
2026 SF
2025 OAK
2024 AUS
2023 AUS
2022 AUS
2020 SF
All talks
A
L
L
T
A
L
K
S
Career Panel - Leveling Up in Your Career as a Data Scientist/Engineer
Learn how to level up your career as a data scientist or engineer
Speakers
Nick Chamandy
, Scientific Director, Lyft
Jonathan Lenaghan
, SVP, Product Development & Chief Scientist, PlaceIQ
Josh Schwartz
, Co-founder & CEO, Phaselab
View transcript
00:00
My name is Jared. I'm a developer evangelist at a company called Galvanize. So what we
00:28
do is we build community around data science and web development. We have a data science
00:33
training program at Galvanize and if you have any questions about that program, I'm happy
00:37
to talk about it after our panel. And we're just going to go around, Robin, now and each
00:41
of us are going to introduce ourselves. So go ahead.
00:44
I'm Josh Schwartz. I'm the head of product engineering and data science at Chartbeat.
00:49
How much intro do you want? Do you want the one line or the...
00:51
Give me like a minute.
00:52
Great. All right. I can do a minute. So Chartbeat is an analytics product focused on doing web
00:59
analytics for media companies. So if you're a major media company in the world, a New
01:04
York Times, New York Magazine, CNN, folks like that, you're using Chartbeat every day
01:08
to understand what stories to write about, how those stories are resonating in the world
01:12
and how to promote them around various channels.
01:17
Thank you.
01:19
Hi. My name is Nick Schimandy. I lead the data science team at Lyft. Hopefully, you
01:23
guys know what Lyft is already. There are some coupons that you can have access to if
01:28
you'd like. I studied statistics in grad school and then I worked at Google for eight years
01:33
or so. Hopefully, you all know what Google is as well, so you don't need to go to that
01:37
one. Yes, I learned a ton of stuff there and trying to apply it at Lyft now and make everybody's
01:43
experience better.
01:44
Hi, everybody. I'm Jonathan Linehan. I'm the VP of Science and Technology at PlaceIQ. PlaceIQ
01:49
is a, we're building a geospatial analytics platform. The main vertical that we're principally
01:58
monetizing right now is mobile advertising. I mean, in New York City, it's fairly typical.
02:02
In terms of my background, I have a PhD in physics. I graduated in 2002 and I've worked
02:07
in industry since 2003. I worked in finance for a long time and I joined PlaceIQ about
02:14
five years ago now. I was one of the very early employees there. I think, and one of
02:20
my passions actually is helping people make the transition from principally academics
02:24
into industry. And so, what I found when I was making the transition myself, there really,
02:31
there wasn't a whole lot of support inside the academic community to do that and it's
02:36
sort of a hobby of mine to help people with that. So, I'm very excited to be on this today.
02:43
Thank you guys. Let's give them a round of applause for taking time out of their day
02:47
to be here. So, as I mentioned, we'll have to talk after, by the way, because Galvanize
02:57
trains data scientists, so we should chat about helping people make the transition.
03:03
But as I mentioned, my name is Jared. I'd like to get to know the audience a little
03:08
bit better, so I always like to ask a couple of questions of the audience before we begin
03:12
so we all know who's in the audience who we're addressing today. Who here is a data
03:16
scientist? Raise your hand please. Okay, awesome. Who here is currently a data engineer? Raise
03:24
your hand. Okay. Who is currently finishing up a program, whether it be a bachelor's degree
03:31
or a master's degree in some kind of STEM field? Okay, we have a few in the audience.
03:37
Who here is working full-time but not as a data analyst, or sorry, not as a data scientist
03:41
or as a data engineer? Maybe something else. Awesome. Cool. Nice widespread. Well, we're
03:48
going to jump right in. So I'm going to ask a very basic but very loaded question of our
03:52
panelists. How do you define data scientist versus data engineer?
03:59
It's a great question. So at Charpy, being a data company, we don't have a formal role
04:07
called data engineer. All of our back-end engineers are focused on making data infrastructure.
04:13
We have folks who work on different products, but every product and everybody's work is
04:17
somehow data related. Similarly, we also on the data science side don't have folks who
04:23
are focused on things like visualization because everything that we do as part of the product
04:27
is visualization. So every front-end engineer is thinking about the data vis side. So for
04:33
us, the real question is what's the role of a data scientist versus the role of an
04:38
engineer or for that matter, a product person. What I tend to think of in the analytics world
04:43
is a data scientist is a person who's kind of marrying those three dimensions. So a data
04:49
scientist is the representative of the data in the product. So they're the person who
04:54
is coming up with new ideas and algorithms, ways to use the data, and also sort of on
04:59
the back-end saying, you know, these are systems that we need that either I'm implementing
05:03
myself or I'm, you know, asking folks to implement with me to collect new data, to store it in
05:09
new ways, to, you know, build algorithms on top of it and so on. So, you know, for us,
05:12
it really is the data scientist role is really kind of sitting in the middle. In some ways,
05:17
almost similar to what a product person's role is with a very different set of concerns.
05:24
I can sort of describe how it works at Lyft which is probably not the only functioning
05:28
model but, and Prashant who's our head of data platform is in the audience so he'll
05:31
throw something at me if I get this wrong. But the way I kind of see it is that the data
05:36
engineers are mainly responsible for like the reliability, correctness, and timeliness
05:42
of the data. The data scientists are more responsible for taking that data and using
05:46
it to solve business problems. And then there's some shared responsibilities. So I think like
05:51
the kind of data model and the semantics of the data structures are, to me, a shared responsibility
05:56
because you can't build that in a silo on either side because it usually doesn't work.
06:02
So that's sort of how it works at Lyft.
06:04
Yes. So I would say that at PlaceIQ, I mean, it's very similar. I mean, in the sense that
06:13
at the level of its title, there is a distinction but in terms of the skill set, I mean, it's
06:16
a very broad spectrum. So we have maybe one or two engineers who really focus on front
06:23
end but largely, I would say the rest of the engineering team, probably about 20 to 25
06:30
people, I mean, I think I would characterize them as, you know, certainly data engineers.
06:35
And then I would say that about one third of them could sub in at any other company
06:40
as a data scientist. And on the data science side, I would say about half of them could
06:44
also sub in at any other company as a data engineer. And so, I mean, at PlaceIQ, we fully
06:50
expect the data science team to be able, with at least some help, to push code into production.
06:57
And so, I think it really is a spectrum and it depends on the problem that we're solving.
07:02
Now, we're starting to get into the license, into the business of licensing software as
07:08
well. And so there, there are a whole another set of skills that are required to enter the
07:13
enterprise software business. And there, we do have people who are, I mean, I think
07:17
very dedicated traditional full psych engineers that I don't think you could sub into the
07:21
role of a data scientist but that's new for us, I would say.
07:27
Thank you. For the attendees in the audience who are currently, we had kind of a wide distribution
07:32
of hand raised earlier. For the attendees who are looking to transition to becoming
07:37
a data scientist or to becoming a data engineer, what advice would you give them?
07:46
So from my perspective, I think the fundamental skill that's overlooked very early in the
07:50
process is having a very solid engineering background. And that, and I don't think that
07:56
means that you, that you necessarily, that you have to be a very good engineer.
07:59
really need to have the skills that a full-stack engineer would have five years on the job,
08:05
but having familiarity with the software development life cycle, being able to speak cogently about
08:13
the best practices in Python development or Java development I think is extremely important.
08:18
And, I mean, when I talk to people particularly in fields like physics who want to make a
08:23
transition, they, this is something that I think is overlooked tremendously by them when
08:29
they're looking for jobs. And I think this is usually because they spend six or seven
08:35
years writing code, but they don't really have a good sense for what modern best practices
08:41
are in software engineering. So I do think that the data science training programs are
08:48
extremely valuable, but I think almost as valuable in any way that you can, you know,
08:56
become familiar with modern software engineering practice.
09:00
Yeah, I couldn't agree more with that. I mean, you know, we expect on the data science side,
09:05
you know, a data scientist within, you know, six months of working at Terpy should be pushing
09:10
code to production on the regular and, you know, able to, you know, dig into a database,
09:15
you know, design a new database, you know, dig into API servers, design new API servers,
09:19
and so on. And so that process of learning is often kind of a shock, especially when
09:25
somebody's coming straight out of a PhD program or another academic program. And sort of readying
09:31
yourself, you know, first of all, understanding that that's what the job is, right? And I think
09:35
you're hearing that in more and more places. That really is, you know, being able to ship
09:38
your own stuff is really important. So kind of preparing for that both, you know, sort of
09:44
mentally and also having the skills is a huge thing. You know, I think you're, you
09:52
know, PhD in physics, having worked on a lot of sort of messy scripts is a great example
09:57
that probably every data scientist hiring has dealt with where you have somebody who's
10:02
probably developed a lot of bad habits and they come in and you can often do, you know,
10:07
sort of more poorly in an interview process than you actually should because you've developed,
10:13
you know, sort of quirky software engineering habits by working on, you know, code that
10:17
there's no reviewer on that after you publish your paper, nobody's ever going to look at
10:21
again and so on.
10:22
Yeah, I think I'd add a skill that's kind of complementary to the ones you guys discussed
10:26
and that is sort of getting your hands dirty with data, with real data, and trying to answer
10:32
questions that are actually really relevant. So there are tons of, you know, free data
10:37
sets out there. There are competitions. There are many things you can dive into. So if you
10:41
have the time, I think that's a great way to learn how to problem solve and in our interviews,
10:46
you know, we test that problem solving ability. How can you, so there's a term that Diane
10:50
Lambert from Google coined recently which is, how do you think with data? And I think
10:54
that's a skill that we really look for and you only sort of acquire that skill by just
10:59
getting your hands dirty.
11:00
That's a really good point, Nick. So what are some of your favorite ways for the audience
11:06
to practice and get that experience? So you mentioned Kaggle. There's a lot of open source
11:10
projects out there. Do you have any particular favorite projects that you recommend people
11:13
get involved with? I was just chatting with Continuum Analytics in Austin last week and
11:18
they were telling me about some of their favorite open source projects. I'm curious, like, what
11:22
are yours and what do you recommend the audience check out to get more software engineering
11:26
experience?
11:28
So in terms of data sets, one that we use a lot, and this is a very transportation specific
11:32
one, is the New York City taxi data. So that's a very rich set of data that we like to play
11:37
around with.
11:38
Awesome. Do you guys have any recommendations?
11:40
Yes. So, I mean, just because of the domain that we're in, I mean, we use the New York
11:48
City set a lot for, I mean, for small projects when you're late stage in the interview process.
11:56
And I think the importance of getting your hands dirty with data is that, I mean, it's
12:04
extremely important that you develop a facility or intuition about how to work with data.
12:09
And so, and this is something that only comes with a lot of practice. And so in the same
12:18
way that if you're in a physics Ph.D. program, you tend to develop a sort of intuition, physical
12:29
intuition that you aren't necessarily, that you don't necessarily have when you're born.
12:37
But in the same way, just doing, you know, small experiments with real data, there's
12:45
no substitute for that. And so, I mean, just in terms of getting access to interesting
12:53
sets of data that solve real world problems, the Kaggle competitions, I don't think you
12:59
can beat that.
13:00
And by the way, you can start small. I mean, there's like data sets in R that you can just
13:04
play around with first and then sort of keep challenging yourself with bigger and bigger
13:07
and more challenging problems.
13:09
I also think, you know, for folks who are sort of ready to dive into problems, you know,
13:14
a great thing that I think is also a good for the world thing to plug is there's an
13:18
organization called DataKind that, you know, does data science volunteering. You work with,
13:22
you know, a set of non-profits and sort of volunteer your time. It's a great opportunity
13:26
to both do, you know, important work in the world and also, you know, pick up skills and,
13:31
you know, the set of people who work on DataKind projects is a pretty awesome group. So it's
13:35
a great way to meet folks and understand what's happening in the industry.
13:40
So, Jonathan, you mentioned data science programs earlier. What are the advantages and disadvantages
13:47
that you three see in students who are coming out of data science programs, whether it be
13:52
a three-month boot camp or whether it be a year-long master's degree? What are some of
13:57
the advantages and disadvantages that you see? And then when you recognize potential
14:02
and you decide to pull the trigger and hire one of those candidates, how do you help them
14:05
overcome the disadvantages?
14:09
So I think as an advantage of, you know, thinking specifically about, you know, sort of accelerator
14:16
programs that take people often from academia and into data science, an advantage is that
14:21
there's often a small, like we were talking about with engineering knowledge, there's
14:24
often a small delta between the set of stuff that a new grad PhD knows and the set of stuff
14:30
that will get them hired at the job that they want. And often you don't really know how
14:34
to talk about your work. You don't really know how to do a coding interview and things
14:38
like that that will block you from ending up where you want to be. But with two or three
14:42
months, you can be positioned to be, you know, to just be a much, much, much stronger
14:46
candidate. And I think those programs can pay huge dividends. And also in terms of connections,
14:54
you know, we've hired folks out of those programs who we just never would have come across and
14:58
would have never come across us otherwise.
15:02
The one fear that I have with some of these programs is they often can be, can lead people
15:08
to be very buzz wordy, right? So I'll see folks, you know, whose resume says, you know,
15:13
I am an expert in Scikit-Learn. I'm an expert in MongoDB. I'm an expert in, you know, Redshift.
15:17
I'm an expert in Spark. And you're like, are you really, you know, did you, did you, are
15:21
you actually an expert in those or did you, you know, sort of import SKLearn, run a couple
15:26
functions and so on? And so I think the key is that, you know, when we're interviewing
15:31
folks, we're looking for expertise, right? So more than a long list of things that you've
15:35
dabbled in, I want to see that you're really good at something. If that something is physics,
15:39
it's fine, right? But I want to know that you really know that stuff. And so I think
15:44
in thinking about those programs, think about them as ways to learn the industry and ways
15:48
to learn, you know, a bit of software. But I think, you know, make sure that you're not
15:54
just sort of thinking about a list of small skills to pick up.
15:57
Yeah, I think, you know,
15:59
along similar lines, it's kind of a bit of a breadth versus depth argument. I think one
16:05
positive thing is that, and this is related to what you said, that because these things
16:11
are relatively short, you end up learning about the most kind of popular methods that
16:17
startups and small tech companies are using today, which is I think if you go through
16:21
the sort of PhD and then go work at a big company for a number of years like I did,
16:25
you sort of miss out on that. You're not as up to date on sort of the newer technologies,
16:29
you know, like in grad school, I used R and MATLAB and then at Google, they have their
16:34
own way of doing things with a very limited set of different tools. And when I left Google,
16:38
I discovered all these other things that were out there that sort of had been built in-house
16:42
at Google in a very different way. And so, I think it's basically exposure to those things
16:46
that you're getting in some of these boot camp type programs. But again, I think the
16:51
drawback is you may not have, depending of course on your background. So, I think the
16:55
drawbacks depend on what you did before. But you may not have as much of a depth in
17:00
terms of the underlying math behind some of the topics you're learning about. And so,
17:05
that would be, you know, the one, that's the one concern we sometimes get with candidates
17:10
coming out of those programs.
17:11
Yeah, I would say that, so I'm a very large advocate of the programs themselves. And I
17:17
think so, and to echo what Josh said, I mean, there really is a, for people who go into
17:22
these programs, there really is a very small delta between the skills that you already
17:29
have and the skills that will get you through the interview process. And so, I mean, so
17:35
if you take a very sort of jaundiced view of the interview process, I mean, the people
17:39
who are, no matter how, even with the best of intentions, the interview process is set
17:46
up to screen out people as quickly as possible. So, if there's even a single thing on your
17:51
resume that, you know, somebody finds annoying, you're already at a very large disadvantage.
17:58
And so, I think the programs are extremely useful for networking, for getting you in
18:05
front of people, for helping you with the interview process. And it definitely level
18:12
sets these sort of skills that you need to pass interview. And so, I mean, I think one
18:22
drawback, I would say, is people come out of the program, the resumes, aside from what
18:28
you did before, they tend to be a bit more undifferentiated in the sense that, like,
18:35
everybody writes down Scikit-learn and Random Forest. And so, that's not a real differentiator.
18:43
But, so I would also say one, another reason why I like the programs is that people typically
18:49
enter them are extremely motivated to make a switch in life. And so, they definitely
18:54
want to go into a different profession. And so, I mean, that already, I think, is a huge
19:01
plus when you're interviewing somebody.
19:03
Yeah. Just to add on to what you said, I was sort of laughing at myself as you talked about
19:07
the undifferentiation. You know, when you're hiring, it's a very different perspective
19:12
than when you're job seeking. And one funny thing about interacting with folks coming
19:17
out of these, out of programs, is that there's a big batch of folks coming out of the program
19:20
at the same time, which means that you're looking at a stack of 30, you know, PhDs from
19:26
elite universities, master degrees from elite universities, who all then list the same,
19:30
you know, 10 data science skills, and are all super motivated people who decided to
19:35
take the time and go through this program. And it makes hiring really hard, right? And
19:40
it makes it really, you're actually in, ironically, by going through the program, you've put yourself
19:45
in a more difficult position than if you applied one-off. Almost anybody who goes through one
19:49
of these programs, if they just emailed in their resume, I would be like, wow, this is
19:53
awesome, we've got to get this person in today. But when you're in that stack of 30, I feel
19:56
like I have to call down to five, and that makes it difficult. So I think, you know,
20:01
thinking about, you know, how you're pitching yourself coming out of it, and rather than
20:05
just sort of relying on the program, picking, you know, which companies that are partnering
20:09
with the program that would I actually be passionate about working at, and like making
20:14
that very, very, very clear to those companies is really important.
20:18
I'm going to go out of order. I have this sheet here that's kind of my, it's helping
20:24
me out in asking these questions, but I was going to go in a particular order, and I have
20:29
to deviate now because you said some really awesome stuff. And to follow up on what you
20:33
said, how do I decide which path to take in terms of data science versus data engineering?
20:40
And then let's say I'm already committed to a path. We have a bunch of people in the room
20:43
who are already data scientists or already data engineers, or maybe they're recent graduates
20:47
from a master's program or a boot camp. How do I decide what to specialize in? So how do
20:53
I differentiate myself? So am I really good at visualization, storytelling? Am I a data
20:59
scientist that is an excellent programmer? Am I an excellent predictive modeler? What's
21:04
my secret sauce, and how do I choose that? How do I find it?
21:12
That's some profound life advice. You know, to me the paths are so different that it's
21:18
almost, I think it's probably obvious to a person considering the paths, right? So
21:25
being a data scientist like being any other scientist is profoundly frustrating and depressing
21:29
in many ways, right? Like most things that you do won't work, right? Like most, you have
21:34
to be, just like being in grad school, you try things, they don't work. You try other
21:39
things, they don't work. You spend a month on a project, it goes nowhere. Nobody likes
21:42
your algorithm. You can't figure out how to do what you were trying to do. And that is
21:48
an extremely rewarding and stimulating process, but is like doing science, right? And engineering,
21:56
I think, you know, I happen to be really excited about both, but engineering actually, the
22:01
problems are hard, but you sort of have this moment where you're like, I know I can solve
22:04
this. And you get that, you know, that payoff where, you know, your code compiles and you're
22:08
like, woohoo. And that doesn't happen in, you know, in data science. So I think, you
22:14
know, which of those spectrums you live on, I think, is probably clear, but to me, that's
22:19
the real difference.
22:20
Yeah, I would agree with that. I think even within data science, which is what I can speak
22:26
to more, it's a little bit unnatural to say like, hey, I want to be a data scientist.
22:30
Now, let me figure out what I want to specialize in. I think it usually happens the other way
22:34
around. You, you know, you always love math and then you say, hey, I can apply this and
22:38
become a data scientist. Or you always liked, you know, writing algorithms to optimize things
22:43
or whatever it might be. So I think of data science is not a perfect analogy, but a little
22:47
bit like social science where you don't just say like, oh, I want to be a social scientist.
22:51
Now, what should I work on? You actually say, oh, I'm passionate about linguistics or sociology
22:55
and then later on, you identify yourself as a social scientist.
22:59
Yes, I mean, so, I mean, so, I mean, I myself am, I mean, I'm very drawn to both of them,
23:06
but I mean, just looking at the people in my company and people I've talked to, you
23:10
really want to think about, you know, what causes you that moment of elation. Because
23:13
at some point, even if your job is frustrating and you hate it, there, it is going to have
23:17
to be, I'm sorry, not that, not that you hate it, but there are times that you will. I mean,
23:27
you, but in order for it to be rewarding, you have to have those moments of elation.
23:34
And so, if discovering something interesting, I mean, really gives you that feeling, then
23:44
I would stray towards the data science path. If what you're interested in is building large
23:50
systems and you're extremely passionate about having something of great complexity running
23:56
all the time, 24 hours a day.
23:59
hours a day, and you really like monitoring, which sounds
24:04
crazy, but some people do, then I would go the route of
24:09
the data engineer.
24:13
So I have a question from a speaker that I was speaking to
24:15
earlier today.
24:16
It's Peter Lenz, who was recently in this room, and now
24:18
he's across the way.
24:19
Oh, there he is in the back.
24:21
He finished up his office hours.
24:23
So Peter asked, let's see here, what advice do you have
24:29
for people who do not come from the traditional stats
24:31
and computer science background, but have found
24:33
themselves at a job where they are called a data scientist?
24:38
And to give you some insight onto Peter's background, he
24:40
was, I believe, a geographer, is that right?
24:44
But he taught himself a lot of the skills along the way.
24:48
Yes, I'd say if you find yourself in that position,
24:51
chances are you do have some relevant skills, either.
24:53
And passion is almost a skill in that sense, right?
24:58
And if you're lucky enough to be in an environment where
25:00
you're working with other data scientists, then I think it's
25:03
a great opportunity to learn from them and maybe teach them
25:06
something as well.
25:08
I think that can work really well.
25:12
I would say that if you've already been in the industry
25:15
for a few years and you've found yourself into that role,
25:22
like you said, I think it's because you are a person who
25:27
naturally acquires new skills and is comfortable with
25:33
ambiguity and role, because that's essentially how you
25:36
random walk there.
25:42
I certainly don't see a problem with that.
25:45
And if I'm interviewing a candidate who has five or 10
25:50
years of industry experience but started off as a history
25:54
major, it's completely irrelevant to me.
26:02
If we zoom out and we look at the entire landscape and look
26:06
at data engineering and data science, what skills are you
26:09
currently seeing a shortage of amongst candidates and amongst
26:15
even data scientists and data engineers that work at your
26:16
companies?
26:17
Where are people weak?
26:20
Or where is there a shortage of skill?
26:23
This isn't a hard skill.
26:25
But I think the hardest thing for most data scientists to
26:29
get to is being able to do product thinking.
26:34
So in the end, if you're working at a technology
26:38
company as a data scientist, the output, you may be doing
26:43
internal analysis where the product is your analysis.
26:46
Or you may be working on a product directly.
26:49
But actually, this sounds trite to say, but your job
26:53
is coming up with problems and then solving them.
26:56
And that's very different than in an academic setting.
26:59
And coming up with the sense of taste for what makes an
27:03
interesting problem and also an attractable problem takes a
27:06
really, really, really long time and is fairly
27:11
domain-specific.
27:12
And that's the thing that I think is hardest to get but is
27:16
the most valuable one you have.
27:19
I would add, I think, maybe not a skill that's necessarily
27:24
missing yet, but that's a little bit underrated relative
27:27
to skills that are being taught in a lot of these
27:30
programs, and that is inference.
27:32
So there's a lot of focus now on prediction problems and
27:36
machine learning and black box type stuff, which I think has
27:38
a place in many problems.
27:40
But there's a lot of nuance around even when you're
27:43
framing a problem or coming up with an idea and understanding
27:48
the data, I think causal inference in particular is
27:51
something that's pretty underrated.
27:53
And we don't see a lot of people coming in with exposure
27:56
to that, or if they do, they have very traditional
27:59
economics or statistics backgrounds and maybe
28:01
underpowered in terms of computing and machine
28:03
learning.
28:03
So I think merging those more is going to be a big bet for
28:10
the future.
28:11
And even extending to things like A-B testing, I don't
28:14
know if anybody saw the talk by, I think,
28:16
Sergey was his name.
28:18
We're facing many of these same challenges where
28:22
something that's kind of taken for granted is actually way
28:25
more nuanced than you might think and requires some
28:30
ability to do inference and not simply just power through.
28:34
So I think those are both wonderful answers.
28:37
To those two, I would also add that estimation is a skill
28:42
that I think people need to hone.
28:46
So being comfortable with giving order of magnitude
28:50
estimates for things is, so for people who come from some
28:55
specific fields, this, just through the course of your
28:59
academic work, is sort of natural.
29:02
For other people, just giving an answer to somebody within
29:07
a factor of two or five, they become extremely
29:10
uncomfortable with that.
29:11
And so I would say that you can get very far without having
29:17
to do a whole lot of work in computation in certain
29:20
circumstances.
29:21
Just being comfortable making estimates based upon what you
29:29
already know.
29:29
And so I know for sure a place like you, we've wasted tons
29:34
and tons of computational cycles and storage when a
29:39
simple back-of-the-envelope calculation would have done.
29:43
And so I'd like to elaborate a bit on Josh's point.
29:46
When I came to a place like you, I had no idea what a
29:48
product manager did.
29:50
I had no sense at all.
29:51
I worked in physics.
29:52
I worked in finance.
29:53
And this notion did not exist.
29:55
And so product thinking, and in particular, even if you
29:59
aren't doing a lot of storytelling, being able to
30:02
talk to people on the business side is extremely important.
30:05
And understanding what product management is or product
30:08
development is is an important skill.
30:11
And I think a good test is whether or not you think you
30:17
can explain what a product person does.
30:21
That's interesting.
30:21
Yeah, we try to get a sense for whether the candidate has
30:25
good intuition for the product and what the constraints of
30:30
the business are in terms of coming up with a solution.
30:32
We used to have a product manager on most of our
30:34
interview loops.
30:34
Unfortunately, we had to stop that because he literally
30:36
voted no to every single candidate.
30:41
So we have product thinking and product taste.
30:44
We have inference and estimation.
30:47
How can we, all of us in this room, and we as a data
30:50
community, how can we decrease this shortage?
30:54
And how can we help decrease the shortage of skills?
31:00
It's a really hard question because those
31:02
are all sort of softer.
31:05
Yeah.
31:08
I think for us, one of the keys is having a diverse team.
31:11
And that's along many axes, including academic diversity.
31:14
So for us, it works really well because we have a
31:17
diverse spectrum of problems to solve, from optimization
31:21
to things like inference, prediction, all kinds of
31:24
stuff, and so we just naturally get a lot of
31:26
different people applying and therefore
31:28
different points of view.
31:31
So I think for us, we can teach each other the gaps
31:36
that may be there.
31:42
So I had one other question from Peter.
31:44
I don't know if he's, is Peter still here?
31:46
Is he gone?
31:46
He's outside.
31:48
Maybe I'll ask this one in a bit when he comes back in.
31:51
So another question that I had is.
31:59
As you start to consider promotions internally for your data scientists and data engineers,
32:06
I guess, what are the common signs that somebody on your team is ready for the next level?
32:12
What differentiates a B data scientist or data engineer from the A scientist or A engineer?
32:18
And furthermore, what is that next level?
32:23
You know, for us, I mean, what's true on the data science side is also true on the engineering
32:28
side, that, you know, moving up levels of seniority is about being comfortable dealing
32:32
with different amounts of abstraction, right?
32:35
So, you know, a more junior person, you know, you can sort of give a concrete problem.
32:39
Okay, implement, you know, A-B testing for this thing, and the person can do it.
32:43
You know, they can, you know, code up their relevant math and, you know, and write the
32:48
right code.
32:49
But the problem statement is pretty defined.
32:51
A more senior person, you can give an abstract problem, right?
32:54
We want to solve this thing, and they can decide, okay, A-B testing is the right methodology.
32:57
Are we doing a split test, or are we doing a bandit?
33:00
Things like that.
33:01
And as you, you know, move up the stack more, the person, you know, can define the problem,
33:05
right?
33:06
Oh, I think that we should do this feature, and I think we should A-B test it, and here's
33:08
the A-B testing algorithm.
33:09
So, I mean, that, you know, that sort of moving up the level of abstraction, I think really
33:14
in almost all jobs is what, you know, moving up the stack of seniority looks like.
33:18
For data scientists, what that generally means is that people are more comfortable
33:23
identifying problems.
33:25
So, you know, a new grad data scientist, we might say, you know, implement this specific
33:32
thing, and a very senior data scientist is just going to be, you know, identifying the
33:37
problems that they should work on and then solving them.
33:40
Yeah, I would agree with all those things.
33:43
Also, I mean, for us, I think the number one thing we look at is impact, and that can mean
33:48
different things in different situations.
33:50
So for some people, they just want to be an individual contributor and, like, have their
33:53
head down, and maybe they're developing new methodology that, you know, unlocks big gains,
33:57
so that's one way to have impact.
33:59
Maybe they're becoming a tech lead and sort of magnifying their impact by mentoring others
34:04
and pushing, you know, others to do new things.
34:08
Maybe they're just coming up with new ideas instead of, like, sitting quietly in a corner
34:11
with this idea or telling their three friends about it.
34:13
They're actually championing this idea to the point where it gets, you know, prioritized
34:16
in a roadmap and actually ships.
34:19
So I think there are different ways of measuring impact, and different people maybe tend towards
34:23
different ways of doing it, but that's, for us, the single biggest bucket, I would say.
34:27
Yeah, I would say the only thing that I would add to those is I think the level up happens,
34:34
for me, when someone starts to really gain domain expertise or knowledge in the industry
34:41
that you're working in.
34:43
And so, I mean, just to take myself as an example, so, I mean, I never thought, you
34:48
know, when I was younger that I'd be working in mobile advertising.
34:53
It just seemed like, why would I ever want to do that?
34:57
And it turns out it's something that, you know, I'm very fascinated by now.
35:03
And so, I think when you're making a shift to industry, it's almost always the case that
35:06
you'll be working in an industry that you never thought was interesting before you went
35:11
into it.
35:12
And so, there are very specific things about each industry that you absolutely have to
35:18
know to make impactful products for your company.
35:23
And, you know, just to give a concrete example, so a lot of people, when they first start
35:29
at Place IQ, they think that data is data and the data generation process itself isn't
35:38
that important.
35:39
A lot of the gains that we've actually made inside are really understanding in detail
35:45
how hybrid positioning works in location services on your phone.
35:52
And so, that's very, very specific to domain knowledge that I wouldn't think would have
35:56
anything to do with mobile advertising, but it turns out to be extremely important.
36:04
And so, I think the people who have accelerated very rapidly at my company are those who've
36:12
really taken the time to get a very deep knowledge about the industry and how the data is actually
36:20
generated.
36:21
Thank you.
36:22
The next question is from an attendee from earlier today.
36:28
I forgot their name, but I'll see you later today.
36:33
The question was, how important is it for a data scientist to be able to code and how
36:39
well should they be able to code?
36:42
And then, how difficult is it to teach scientists to engineer?
36:48
I think it's super important.
36:49
I mean, it's importance can't be understated.
36:53
If somebody can't ship code, they can't do their job, at least, you know, for us.
37:00
You know, I don't expect a data scientist interviewing to, you know, understand, you
37:10
know, real heavy-duty software engineering.
37:15
I don't understand them to have, you know, worked with, you know, Kafka or, you know,
37:23
databases, any of the systems that we use for data processing.
37:26
But I certainly expect them to be able to write well-formatted, carefully thought-out,
37:31
algorithmically signed code.
37:33
It's really, really important.
37:35
And they have to like doing it.
37:36
Like, you know, there are times when you talk to folks who sort of view, like, code as a,
37:41
you know, as a thing that you have to do.
37:44
You know, I really like doing pencil and paper math or, you know, lab work, but I can write
37:49
code if I need to.
37:50
And you want somebody who actually is excited about building because, you know, it's kind
37:53
of what the job is.
37:56
Yeah.
37:57
I think, I mean, I agree at Lyft as well, being able to code is necessary.
38:03
There's obviously different levels of that.
38:04
So I think one thing to keep in mind is that data science means different things at different
38:08
companies.
38:09
And so, I don't know that the requirements are, you know, that stringent everywhere.
38:13
For us, I think, like, you know, having familiarity with SQL is important.
38:19
Having like being able to write scripts is obviously really useful.
38:24
And then for most people coming in, like, knowing object-oriented programming at some
38:28
level because they're going to be expected to at least ramp up in that in order to write
38:32
production code.
38:34
So having it from the get-go is obviously a bonus.
38:37
In terms of, like, teaching data scientists how to be engineer or be more engineering-oriented,
38:43
I'm not really an expert on that.
38:45
But I would say that for me, it was, you know, just working closely with engineers really
38:51
helped me.
38:52
I mean, that first code review that I sent to a really senior engineer was, like, slightly
38:56
traumatic experience but also, you know, lessons that I never forgot.
39:00
And so, I think encouraging especially junior data scientists to work closely with experienced
39:06
engineers or other data scientists who are, you know, strong coders is really valuable.
39:09
Yeah.
39:10
I mean, I agree with that.
39:11
I mean, I think it's extremely important.
39:13
There certainly are roles where the engineering aspect is not as important.
39:19
But, I mean, just in terms of finding a job in the industry, I mean, you, it's going to
39:26
be much, much easier if you can write clean, well-formatted code.
39:33
I think in terms of having to know, you know, modern data infrastructure, Kafka, Spark,
39:40
et cetera, if you aren't already working in the industry, I wouldn't spend much time on
39:45
that.
39:46
I mean, it's hard to learn that on your own.
39:49
And that's something you can certainly learn on the job.
39:52
But the ability to write code and more importantly, to like to write code, I think is essential.
39:59
There are roles that won't require that, but they're
40:05
becoming fewer, fewer, and fewer.
40:09
Well, Peter's out there talking, so I'm just going to
40:11
ask his question.
40:12
He can watch the video later.
40:14
So Peter's question was, from Peter Lenz again, there's kind
40:18
of an elephant in the room in our industry.
40:20
We all work in the technology industry, and it tends to be
40:24
slightly ageist.
40:26
Right?
40:28
So here's a scenario.
40:30
I'm a 30 to 35-year-old data scientist.
40:33
Sexiest job of the century.
40:35
I don't know if anybody's benefited from that title,
40:38
like, sexiest job in the world.
40:40
But try it at, like, maybe you'll get a free gym
40:43
membership or a free beer sometime.
40:47
I'm the hotness.
40:48
I'm highly in demand.
40:49
I'm a young gun.
40:50
I'm the hotness.
40:52
The data scientists and the data engineers in the room,
40:53
you're the hotness.
40:57
What do I need to learn or do to be relevant in 10 years?
41:03
And does a data scientist or a data
41:05
engineer have a shelf life?
41:08
Oh, I don't think so.
41:09
And I certainly hope not.
41:10
I don't know.
41:17
I mean, just make cool stuff, right?
41:19
I don't know.
41:19
I mean, I think that I certainly am more excited about
41:28
talking to somebody who's been an engineer for 25 years and
41:31
has been around the block than I am about a new guy.
41:33
So yeah, I mean, if you're doing good work, right now is
41:37
obviously a crazy time for technology in general.
41:40
But if you're building great stuff, I think you don't have
41:45
anything to worry about for a long, long time.
41:48
Yeah, I agree.
41:49
It's a little hard to take that quote seriously if you
41:50
know Hal Varian, who is the one who said it.
41:53
But yeah, I certainly hope that there's no shelf life.
41:57
I think if you've studied these disciplines, you're
42:02
probably someone who is happy to learn new things and evolve
42:06
your thinking over time.
42:07
And I think there's no substitute for experience when
42:10
it comes to the core problems that keep coming up over and
42:13
over, maybe in a slightly different guise.
42:15
But I think the more experience you have, the more
42:18
quick you are to solve new problems, too.
42:21
So I mean, I certainly don't think there's a shelf life if
42:23
you're willing to learn new skills.
42:26
And so if you work in technology, that's something
42:27
that you have to do over and over and over again.
42:29
And so if you have that capability, then I don't think
42:34
there's a problem at all.
42:36
In terms of ageism, I think there may be a small bias, but
42:39
it's not huge.
42:42
I mean, I'm the third.
42:43
My company is like 150 people, and I'm the third oldest
42:46
person there, and I'm only 42.
42:49
And I think if I hadn't started there so early, coming
42:51
in may have been a little awkward, when everybody is
42:54
like 25 and 27.
42:56
But I mean, it does nothing to affect the quality of your
43:03
work at all.
43:05
Thank you, guys.
43:08
Is there a glass ceiling for data scientists or data
43:10
engineers who don't want to manage, or like you mentioned
43:14
earlier, who are not comfortable with those levels
43:16
of abstraction?
43:18
Or who are on a team where there just aren't management
43:20
positions open?
43:22
So I'm on a team of 20.
43:24
We have a director.
43:26
Where do I go from here?
43:29
I think it depends on the company.
43:32
And if you're joining an early stage company, sort of the
43:38
sky's the limit for data scientists, right?
43:40
My own experience as a data scientist, I started at a
43:43
chart beat as a data scientist.
43:46
I thought my job was to sit in a corner and write algorithms.
43:49
And as we went from a data science team of one to two to
43:53
five to nine, I went on to lead that team, went on to
43:57
lead the engineering team, went on to lead the product
43:59
team, like if you're doing impactful stuff and you join
44:03
early, there's a ton of potential to grow.
44:06
At more mature companies, the standard setup is that you have
44:11
individual contributor tracks and manager tracks.
44:15
Certainly, if you're joining a place, you should look for a
44:18
place that has an individual contributor track that
44:21
actually is real.
44:23
But I think there's so much good work to be done in the
44:27
data science world that I don't feel like there's a
44:30
limiter in terms of requiring people to move into management.
44:35
And I actually think, despite having done it, I think in
44:39
many ways, staying an individual contributor for as
44:42
long as possible is a good move.
44:44
Once you move into management and your life becomes
44:46
meetings, you stop learning about hard technical skills,
44:51
at least, and picking up technical skills for as long
44:54
as you can is both way more fun and rewarding and also
45:00
pays off in the long term.
45:02
Yeah, I would agree.
45:03
I think definitely at Lyft and Google and other places,
45:08
there's certainly an individual contributor track
45:10
that goes sky high.
45:13
I would say, even if you can be protected from people
45:16
management, I think one thing you probably can't get away
45:19
from is leadership of some sort.
45:20
And usually, that's technical leadership.
45:22
And I think the reason is that if you are a level 10, on
45:26
whatever scale, data scientist, but not doing any
45:30
technical leadership, then it's kind of a waste of your
45:32
skills because you only have 24 hours in a day to do the
45:34
amazing work you're doing.
45:35
And you can amplify that a lot by mentoring three or four or
45:40
five other people.
45:42
So I think technical leadership, though, is still
45:45
rewarding in kind of the same way that being an individual
45:48
technical contributor is.
45:50
I would agree with that.
45:52
I would say that in terms of executive level roles, I don't
45:57
think that they really exist except in
45:59
the level of the CTO.
46:01
I mean, most chief data scientists are non-executive
46:05
positions, but I don't really think that's what we're
46:08
thinking about here.
46:09
I mean, in terms of salary, I certainly don't think there's
46:14
a glass ceiling.
46:15
And on the technical leadership side, I mean, I
46:20
think we've had very similar tracks where, I mean, when I
46:24
started at Place IQ, I was the only data scientist there.
46:31
I had no interest in doing any sort of management or
46:33
technical leadership.
46:34
And then I took over that team and then became VP of
46:39
engineering.
46:39
Now I'm working in product.
46:40
So I mean, as long as you're willing to learn new skills,
46:43
either on the IC track or in the technical leadership
46:50
track, I really don't think there's anything
46:51
holding you back.
46:52
And in particular, if you're coming from the data side, I
46:56
mean, I think you're at a much greater advantage to take over
47:02
or lead entire components of the product development
47:06
organization, not just engineering or data science.
47:15
Quick pulse check.
47:16
How are we doing on time?
47:17
We're good?
47:18
We've got five minutes?
47:21
OK.
47:22
Let's give our panelists a round of applause again.
47:25
Thank you.
47:31
I'd like to use our last five minutes for Q&A. And I'll be
47:35
the guy who runs around the room and hands out the mic.
47:37
So let's kick it off.
47:41
You first?
47:46
So would you say that there's a limit in those who have not
47:52
gotten a very thorough mathematical
47:54
background from a higher institution, a limit of how
47:57
much data science they can actually understand?
47:59
What do you think can pick that up in production?
48:05
It's such an interesting question, you know, it's a very hard question to answer.
48:11
You know, most successful data scientists have very serious mathematical educations.
48:17
On the other hand, the irony is that most applied data science doesn't involve sophisticated mathematics.
48:25
It involves reasoning about numbers, but not bleeding edge optimization or machine learning algorithms.
48:33
And so it's not at all clear why it's the case.
48:37
There's this family of problems that I sort of call problems that you shouldn't need a PhD to solve, but somehow you do.
48:44
And so there's absolutely no reason why a motivated, creative, quantitatively inclined,
48:50
but not quantitatively educated person couldn't learn those skills.
48:53
But it is, you know, I think being frank, it is hard.
48:57
And most people who do it have, you know, CS degrees, physics degrees, linguistics degrees, you know, neuroscience degrees, and so on.
49:09
Yeah, I think more than anything, I agree with Josh, it's kind of like a signal that the person has the ability to get deep into a quantitative problem.
49:16
But it's not necessarily the thing that causes them to have success.
49:19
Like I was talking with my colleague yesterday, when's the last time we used measure theory in one of our data science problems, right?
49:24
It doesn't really happen.
49:26
But I think it can be a signal, but it's not necessarily a requirement.
49:32
We hired someone recently who, you know, had a sort of a biological sciences PhD, later on did, you know, a galvanized program, I think.
49:42
And then, you know, interviewed extremely well, because she was able to just reason about the problems and the data in a very organized way,
49:50
and was able to come up with reasonable solutions within the business constraints and all those things.
49:54
And so, it's definitely not a prerequisite.
49:58
Yeah, I mean, I think having strong quantitative reasoning skills and intuition is really all that matters.
50:07
But, I mean, I think having a degree like that, I mean, it's a predictor, but, I mean, it's not something that I would dismiss anybody out of hand.
50:20
The one thing that I think, you know, grad school gets a person used to, to, you know, what I mentioned earlier about failure,
50:27
grad school gets a person very used to failure and to dealing with uncertainty, and it's often the biggest hurdle.
50:33
You know, we certainly hired folks straight out of college or out of industry jobs, and getting into data science and getting used to failure and uncertainty
50:43
is often the biggest thing that people need to come to terms with for at least, you know, their first six months or a year.
50:51
Yeah, and I would say there are, like, some basic skills that we look for in candidates that you don't need a PhD in order to possess,
50:58
but things that, you know, basic mathematical skills that you should have, understanding probability, understanding variance, those kinds of things.
51:06
So I think that's important, but those are not things that require, you know, five years of studying.
51:13
Question over here. Okay. I think we've got time for, like, two more. We can squeeze them in.
51:22
So we talked a lot about data scientists and data engineers today. I have a two-part question.
51:28
Instead of talking about the individual contributor or the scientist, how do you find success as a technical lead on a data-centric team?
51:40
And what are those qualities that you need to foster?
51:44
Nat is asking the question as a technical lead on a data-oriented team at Terpene, so I'll let someone else take that.
51:51
So what do you mean by data-oriented team? Is that, like, an engineering team that happens to have data science problems?
51:56
Yeah, it could be. I mean, I can only live a little broad, but maybe perhaps a product that shows data or, probably in your case,
52:04
something that, like, leverages data internally, using it to make decisions or something like that.
52:11
Yeah, so I think a lot of the themes we've talked about, right, so, like, balancing the business constraints with the mathematical framing,
52:17
being able to convince, you know, a product manager that a solution is the right way to go,
52:25
being able to work with the software engineers to make sure the interfaces are such that this thing can actually work,
52:33
and also collaborating with other data scientists, because, I mean, typically for, at Lyft at least,
52:36
there are usually at least two to three data scientists working on a particular team, and so, you know, collaborating,
52:44
playing off each other, iterating on each other's ideas, so having a collaborative spirit
52:49
and a very collegial kind of approach is pretty important.
52:53
Yeah, I mean, I think the only thing I would add is the ability to speak multiple languages,
52:59
and, I mean, it's been my experience that people who are very deep in engineering are incentivized by some class of things,
53:11
and on the data science side, it's another class of things at the extremes.
53:15
And so really understanding what motivates people and incentivizes them, I think, is important.
53:22
That's a mistake that I've made many times in the past.
53:28
I have a question, and that's about balancing research side of a data science and the product part of a data science,
53:37
engineering part of a data science.
53:39
So if you're working at a tech startup and your resources are limited, you're not like Google,
53:44
that you can spin up a thousand machines that run one SVD or matrix factorization in a matter of seconds,
53:51
you would have to come up with new algorithms that are computationally efficient,
53:56
and they're on the cusp of the research.
54:00
So how do you balance that time between putting a lot of research time and actually coming up with something that works?
54:10
Yeah, I'll take it.
54:14
So I think if you're working as a data scientist at a company that makes a product that isn't directly what your research is,
54:24
and it's not a huge company, you're using mostly off-the-shelf stuff, right?
54:28
So you may be applying, you're choosing what algorithms to use, you may be implementing things by themselves,
54:37
but for the most part, you're using off-the-shelf databases and you're using libraries that manage things.
54:45
So I think, especially for folks coming out of academia, it's a bit of a transition to get used to,
54:55
like so much of your work is sort of figuring out which algorithm to apply rather than writing new algorithms,
55:01
and that's often a change.
55:04
In the end, it's the result that matters.
55:08
If you can solve something without coming up with a novel algorithm,
55:12
you've done a better job than if you had to come up with a novel algorithm,
55:16
which is exactly the opposite of incentive of if you're trying to get a paper published,
55:20
and that takes some getting used to.
55:23
I think it's actually a difficult problem,
55:26
and usually requires some amount of back and forth with product owners at a company,
55:31
and of course it depends on the size of the company,
55:33
but even for a company like Lyft, which is not that small anymore,
55:36
we still struggle with this balance,
55:38
and I think it can become very tempting from a product perspective to just make incremental improvements
55:45
and not think about sort of long-term research,
55:48
and so we try to weigh the pros and cons of that.
55:52
One example is we have this simulation infrastructure which we've built,
55:56
which is actually a lot of work to build and maintain and extend.
55:59
for different use cases and it's not like directly improving the product, it's
56:04
only allowing us to better evaluate other things that will improve the
56:08
product. And so, but you know, I feel like our product managers are pretty
56:13
accepting of the fact that this is going to, you know, take some of our time and
56:17
maybe, you know, mean that a data scientist is unavailable for two weeks
56:20
while that person is building out this simulation environment, but in the long
56:24
run it's going to pay dividends. So I think it's about partly about convincing
56:28
the product owners that this stuff is valuable and also about just sort of
56:32
believing in it. And I think once the product owner is convinced, I mean,
56:41
it's something that I always put a finite time box around and so
56:48
and that's always like, you know, struggle messaging that in a way that, you
56:54
know, doesn't upset somebody, but that's management.