My Projects
Home
Free AI CV Maker
Data Scientist CV
Elena
Vasquez
Senior
Data
Scientist
elena
.
vasquez
@
email
.
com
+1 (415) 827-4613
San
Francisco
,
CA
linkedin
.
com
/
in
/
elenavasquez
·
github
.
com
/
evasquez
Data
scientist
with
6+
years
of
experience
building
production
ML
systems
in
high
-
growth
environments
.
Specialise
in
NLP
,
recommendation
systems
,
and
predictive
modelling
.
Passionate
about
translating
complex
data
into
actionable
business
outcomes
and
leading
cross
-
functional
teams
toward
measurable
impact
.
1 / 3
SKILLS
Languages
&
Libraries
Python
R
SQL
PyTorch
TensorFlow
scikit
-
learn
XGBoost
Pandas
NumPy
Spark
MLOps
&
Infrastructure
Docker
Kubernetes
AWS
SageMaker
MLflow
Airflow
Git
CI
/
CD
Specialisations
NLP
Recommender
Systems
Time
Series
A
/
B
Testing
Bayesian
Statistics
Deep
Learning
EDUCATION
Stanford
University
M
.
S
.
in
Computer
Science
—
AI
Specialisation
2017 – 2019
UC
Berkeley
B
.
S
.
in
Statistics
&
Computer
Science
2013 – 2017
LANGUAGES
English
—
Native
Spanish
—
Native
Mandarin
—
Professional
working
proficiency
CERTIFICATIONS
AWS
Certified
Machine
Learning
–
Specialty
(2023)
Deep
Learning
Specialisation
deeplearning
.
ai
(2020)
TensorFlow
Developer
Certificate
(2021)
EXPERIENCE
Lumina
Health
—
Senior
Data
Scientist
Apr
2022 –
Present
Vercor
Analytics
—
Data
Scientist
Sep
2019 –
Mar
2022
BrightLine
Research
—
Junior
Data
Scientist
Jun
2018 –
Aug
2019
SELECTED
PROJECTS
Cross
-
Lingual
Document
Summariser
PyTorch
·
HuggingFace
·
FastAPI
Fine
-
tuned
a
multilingual
T
5
model
on
a
custom
corpus
of
50
k
English
–
Spanish
paired
documents
.
Deployed
as
a
REST
API
serving
10
k
+
requests
per
day
with
sub
-
second
latency
.
Market
Basket
Analysis
Engine
Spark
·
GraphX
·
PostgreSQL
Built
a
scalable
frequent
-
itemset
mining
system
processing
12
M
weekly
transactions
.
Generated
product
affinity
scores
used
by
the
merchandising
team
to
optimise
shelf
layouts
,
contributing
to
a
6%
lift
in
cross
-
category
sales
.
Designed
and
deployed
a
real
-
time
clinical
risk
prediction
pipeline
using
gradient
-
boosted
trees
,
reducing
ICU
readmission
rates
by
18%
across
12
hospital
partners
.
—
Led
development
of
a
transformer
-
based
NLP
system
for
extracting
structured
patient
histories
from
unstructured
clinical
notes
,
achieving
93%
accuracy
on
a
held
-
out
test
set
.
—
Built
a
Bayesian
A
/
B
testing
framework
that
enabled
product
teams
to
make
statistically
sound
rollout
decisions
,
accelerating
experiment
velocity
by
40%.
—
Mentored
a
team
of
4
data
scientists
and
established
code
review
practices
,
model
documentation
standards
,
and
reproducible
experiment
workflows
.
—
Built
a
multi
-
armed
bandit
recommendation
engine
for
a
SaaS
platform
serving
2.3
M
monthly
active
users
,
increasing
content
engagement
by
27%
and
session
duration
by
14%.
—
Developed
a
churn
prediction
model
using
XGBoost
and
engineered
features
from
behavioural
sequences
,
achieving
precision
of
0.82
at
recall
0.65
and
driving
a
22%
reduction
in
monthly
churn
.
—
Architected
an
automated
feature
engineering
pipeline
on
AWS
that
reduced
model
iteration
time
from
weeks
to
less
than
2
days
.
—
Partnered
with
product
managers
and
engineers
to
design
and
analyse
30+
controlled
experiments
,
influencing
the
product
roadmap
with
data
-
backed
recommendations
.
—
Developed
time
-
series
anomaly
detection
models
for
IoT
sensor
data
from
industrial
equipment
,
achieving
a
false
positive
rate
below
2%
on
production
streams
.
—
Implemented
a
real
-
time
dashboard
using
Python
,
Apache
Kafka
,
and
PostgreSQL
to
visualise
equipment
health
metrics
across
200+
monitored
assets
.
—
Conducted
statistical
analyses
to
validate
sensor
calibration
protocols
,
contributing
to
a
15%
reduction
in
maintenance
costs
.
—
2 / 3
Drift
Detection
Monitor
(
Open
Source
)
Python
·
Evidently
·
MLflow
Created
an
open
-
source
library
for
real
-
time
drift
detection
in
production
ML
pipelines
using
statistical
distance
metrics
and
adaptive
thresholds
.
Currently
used
by
3
early
-
adopter
teams
in
production
.
3 / 3
More Examples