Aayush Yadav, Senior Data Engineer

San Jose, CA  /  Senior Data Engineer

Aayush Yadav

Data engineer who likes numbers that behave

I build the pipelines that carry a company’s numbers from wherever they start to wherever people need them, on time and correct. Seven years of it, across banking, telecom, and healthcare.

80M
records moved a day
Bank of America
99.8%
pipeline uptime
holding steady
$500K+
cloud spend cut
across three companies
7+
years doing this
and still curious
Scroll
About

A short version of me

Hello, I’m Aayush. I have 7+ years of experience turning messy data into something a business can actually use. Right now I lead a team of 4 engineers at Bank of America.

My day is mostly this: something upstream changes, a number looks wrong, and someone needs an answer before their meeting. I like that work. Finding the break, fixing the cause instead of the symptom, and making sure it does not come back.

The part I care about most is handing people their own data. When analysts can pull their own reports, they stop waiting on me and start asking better questions. That is the whole job, really.

Work

Where I have been

Seven years across banking, telecom, and healthcare. Same job underneath: make the data land on time and make it right.

Bank of America Now

Senior Data Engineer

Mar 2024 – Present Charlotte, NC (Remote)

Lead a team of 4 who own the AWS pipelines behind risk, compliance, and regulatory reporting. About 80 million transaction records move through them every day.

80M records a day
$220K saved a year
99.6% data quality
What I did there Show less
  • Rebuilt the ingestion layer on Glue, Lambda, and Step Functions to onboard 25+ internal and vendor feeds. Reporting data that used to land next day now lands within minutes, and about 20 hours a week of manual file handling went away.
  • Designed the Redshift and PostgreSQL models behind risk analytics. Queries on the 8 TB warehouse come back in under a second, which kept BCBS 239 reporting on schedule through two exam cycles.
  • Rolled out dbt with 200+ automated tests across 4 lines of business. Data quality scores went from 94% to 99.6%, and roughly 100 analysts now build their own reports.
  • Tuned the heaviest Spark jobs on EMR with better partitioning, broadcast joins, and adaptive query execution. Batch runtimes dropped by half and compute spend fell about $220K a year.
  • Moved all 35 production pipelines to Terraform with Jenkins and CodePipeline CI/CD. Deployments that took 2 days now ship in an afternoon.
  • Built CloudWatch monitoring with anomaly detection. Uptime holds at 99.8% and failures get caught in 15 minutes instead of 2 hours.
  • Mentor 3 junior engineers. Both of last year’s new hires were shipping to production within their first month.
PythonSQLSparkdbtAWS GlueEMRRedshiftStep FunctionsTerraformJenkins

T-Mobile

Data Engineer

Jun 2021 – Feb 2024 Overland Park, KS (Remote)

Built streaming pipelines on Kafka and Spark in Databricks for network telemetry and customer usage, peaking around 900 million events a day.

900M events a day
$3M revenue protected
$180K saved a year
What I did there Show less
  • Curated Delta Lake tables ready within minutes of the events landing.
  • Owned billing and customer activity ELT into Snowflake with dbt, around 150 tested models. Marketing and retention went from weekly extracts to same-day data.
  • Partnered with data science on churn feature datasets covering usage, network quality, and billing history. The retention campaigns built on them were credited with protecting about $3M in annual revenue.
  • Migrated 20+ legacy Informatica jobs to Databricks and Delta Lake on S3, cutting job failures by two thirds and saving about $180K a year.
  • Introduced a schema registry and data contracts for upstream producers, which ended a recurring class of breakages caused by unannounced payload changes.
  • Carried on-call and pushed uptime from 97% to 99.5% with retry logic, dead letter queues, and alerting that pages on real symptoms instead of noise.
PythonKafkaSpark StreamingDatabricksDelta LakeSnowflakedbtAirflowTerraform

Cardinal Health

Data Engineer

Apr 2019 – May 2021 Dublin, OH

Built Azure pipelines integrating 20+ sources across pharmacy orders, distribution, and claims, cutting processing latency by 60% for daily operational reporting.

60% less latency
Clean HIPAA and SOX audits
25 hrs saved a week
What I did there Show less
  • Wrote and tuned complex SQL over multi-terabyte datasets. Indexing and partition pruning cut the nightly reporting window from 6 hours to under 3.
  • Automated ingestion into Synapse that replaced roughly 25 hours a week of manual ETL and moved data freshness from T+2 to same day for 8 business units.
  • Developed Python automation and internal APIs for reporting workflows, retiring 3 manual reconciliation processes for the finance team.
  • Implemented retention and governance policies in Collibra for HIPAA and SOX. Every audit across two years closed with zero findings.
  • Worked with BI and supply chain analysts to deliver analytics-ready datasets, taking ad hoc requests from a 5-day queue to same-day turnaround.
PythonSQL ServerOracleAzure Data FactoryDatabricksSynapseADLS Gen2CollibraPower BI
Projects

Things I built because I wanted to

Side work, mostly. Both of these started as a problem that annoyed me enough to fix properly.

Real-Time CDC Lakehouse

60s source to query

An end-to-end change data capture platform. PostgreSQL and MySQL changes stream through Debezium and Kafka into Iceberg tables on S3, with exactly-once writes, automatic schema evolution, and time travel queries.

  • Query-ready in Athena within 60 seconds of the source transaction
  • Fully containerized, deployed with Terraform, monitored with Grafana
DebeziumKafkaSparkApache IcebergAWSTerraform

AI Data Catalog Assistant

~50% fewer repeat questions

A natural language layer over a dbt project and its lineage graph. Analysts ask things like "where does customer_ltv come from and is it fresh today" and get the model chain, test status, and freshness back in plain English.

  • Embeddings over model documentation and metadata with retrieval augmented generation
  • Cut repetitive lineage and definition questions to the data team roughly in half
Claude APIVector searchdbtPython
How I work

Four things I do every time

01

Find the real question

Most data requests are a guess at what someone actually needs. I ask first, so the thing I build gets used.

02

Make it boring

Clear names, tested models, small pieces. A pipeline nobody has to think about is the goal.

03

Watch it in the wild

Alerts on real symptoms, not noise. If something breaks I want to know before the person reading the report does.

04

Hand over the keys

Documentation and self serve access, so the team stops waiting on me for numbers they can pull themselves.

Skills

What I work with

The tools I reach for most are near the top of each list.

Languages

05
PythonSQL (T-SQL, PL/SQL, Spark SQL)PySparkScalaJava

Data engineering

08
Apache SparkApache KafkaApache AirflowdbtSnowflakeDatabricksDelta LakeApache Iceberg

AWS

13
GlueEMRRedshiftS3LambdaStep FunctionsKinesisAthenaLake FormationDynamoDBCloudWatchIAMCloudFormation

Azure

05
Data FactoryDatabricksSynapse AnalyticsADLS Gen2Azure Monitor

Databases

05
PostgreSQLOracleSQL ServerMySQLMongoDB

DevOps and tooling

07
TerraformDockerKubernetesJenkinsAWS CodePipelineGitCI/CD

BI and governance

08
Power BITableauAmazon QuickSightCollibraDimensional and lakehouse modelingHIPAASOXBCBS 239
Education and certifications

The paper trail

Bachelor of Science in Computer Science

Minor in Computer Information Technology

Northern Kentucky University Highland Heights, KY

AWS Certified Data Engineer

Associate (DEA-C01)

AWS Certified Cloud Practitioner

CLF-C02

Microsoft Certified

Azure Data Fundamentals (DP-900)

Contact

Let us talk

If you are hiring for a senior data role, or you have a pipeline that keeps breaking at 3am, send me a note. I read everything and I reply.

Prefer email? Aayush.Yadav1240@gmail.com