YD
YASH DANTALEDATA ENGINEERING / SYSTEMS
HELLO, I'M

YashDantale

DATA ENGINEER

Building scalable cloud, data and infrastructure systems — turning complex engineering problems into reliable, useful products.

Currently focused on
AWSDatabricksAtaccama Data QualityData Engineering
03+Years of Experience
15+Projects & Deliverables
SCROLL TO EXPLORE
01 ABOUT ME

I like building things that make complicated systems feel simple.

I'm a Data Engineer working across cloud infrastructure, data platforms and data quality. My work sits at the intersection of reliable infrastructure and useful data — from AWS foundations and lakehouse architecture to governed pipelines, MDM and AI-ready systems.

I started with Electronics & Telecommunication Engineering and an Honors focus in AI/ML, then moved into enterprise data engineering. Today I'm focused on building systems that are scalable, observable and practical to operate.

03+years in data engineering
9.34engineering CGPA
AI/MLhonors foundation

Cloud

AWSTerraformDockerInfrastructure
01

Data

DatabricksPySparkSQLData Quality
02

Development

PythonJavaScriptReactAutomation
03
HOW I APPROACH ENGINEERING

Reliable by design. Practical in execution.

I care about the part that happens after the demo: maintainability, data correctness, automation, observability and clear ownership.

01

Build for reliability

Design cloud and data workflows with predictable failure handling, clear boundaries and operational simplicity.

02

Automate the repeatable

Use infrastructure-as-code, reusable patterns and automation to remove unnecessary manual work.

03

Treat data as a product

Prioritize quality, lineage, governance and consistency so downstream teams can trust the data.

04

Optimize what matters

Balance performance, scale and cost — especially across distributed processing and large data volumes.

WHAT I WORK ON

Cloud foundations, data platforms and trustworthy data.

My experience spans the infrastructure underneath data systems as well as the engineering patterns that make those systems dependable in production.

01

CLOUD INFRASTRUCTURE

AWS, Terraform and infrastructure patterns for enterprise data-quality workloads, with Databricks in the delivery stack.

02

DATA PLATFORMS

Lakehouse architecture, Apache Iceberg, S3, PySpark and scalable data processing.

03

DATA QUALITY

Ataccama Data Quality, MDM, Golden Record workflows and governed enterprise data.

04

ENGINEERING

Python, PySpark, SQL, APIs, ETL pipelines and automation focused on maintainable systems.

AI / ML IN PRACTICE

Curious about AI. Grounded in engineering.

I started with an Honors foundation in AI/ML and keep building on it through hands-on projects, modern AI tooling and practical data-engineering use cases.

01

Foundation

Honors in Artificial Intelligence & Machine Learning, backed by an engineering mindset around data, systems and experimentation.

AI / MLComputer VisionApplied Computing
02

Applied projects

Built an Image Fourier Transform application and explored LLM / RAG workflows — learning by turning AI concepts into working software and measurable engineering patterns.

PythonNumPyStreamlitLLMs
03
RAGLLMDATA

Keeping up with the stack

I use AI as an engineering layer: connect reliable data to retrieval, LLM APIs and evaluation rather than treating AI as a black box. I stay current through certifications, documentation and small experiments before bringing a new technique into a project.

LangChainOpenAI APIsRAGVector DBs
03+Years Experience
10TB+Data Migrated
10M+Records / Day
99.9%Data Uptime
02 EXPERIENCE

My Journey

A timeline of my professional journey, learning, growing and building impactful data and cloud solutions.

View Full Resume
MAY 2026 — PRESENT

Data Engineer 2

Deloitte

Building AWS infrastructure for Ataccama Data Quality for a leading U.S. banking organization, using AWS and Databricks across cloud infrastructure, data engineering, data quality and enterprise data governance.

  • AWS infrastructure and cloud foundations
  • Databricks data engineering workflows
  • Ataccama Data Quality + governance
CURRENT
NOV 2022 — MAY 2026

Data Engineer

Cognizant Technology Solutions

Engineered scalable lakehouse platforms, ingestion frameworks, MDM/Golden Record solutions, APIs and governed ETL pipelines across regulated pharmaceutical data environments.

  • 10TB+ legacy data migration to Apache Iceberg on AWS S3
  • 10M+ records/day ingestion with 99.9% data uptime
  • 30% reduction in Spark pipeline execution time and cloud compute cost
  • 40% reduction in clinical-trial reporting lag through API integrations
  • 15+ hours/week of manual setup eliminated through Python automation
  • 25% reduction in data-quality errors through programmatic validation
2018 — 2022

Bachelor of Engineering

Modern College of Engineering · SPPU

Electronics & Telecommunications Engineering with Honors in AI/ML. CGPA 9.34/10.

  • Honors in Artificial Intelligence & Machine Learning
  • Foundation in engineering, systems and applied AI
CREDENTIALS · 04

Learning that stays close to the work.

Certifications across data engineering, Databricks, cloud fundamentals and Generative AI — aligned with the systems I build in production.

01DATABRICKS

Databricks Certified Data Engineer Associate

Data Engineering

Certified
02GENAI

Databricks Certified Generative AI Engineer Associate

Generative AI · RAG · LLM applications

Certified
03AZURE

Microsoft Certified: Azure Data Fundamentals

Cloud + Data

Certified
04AI / ML

GSI NVIDIA Technologies Curriculum

AI / ML Foundations

Curriculum
03 SELECTED WORK

Things I've Built

A mix of production engineering experience and hands-on projects — from enterprise lakehouse systems to developer tools and applied computing.

View GitHub
ENGINEERING IMPACT

Numbers behind the systems.

Scale, performance and reliability improvements from end-to-end data engineering work — from ingestion and migration to validation and APIs.

10TB+

legacy data migrated

Apache Iceberg · AWS S3
10M+

records processed daily

enterprise ingestion
30%

pipeline execution improvement

Spark optimization
40%

reporting lag reduced

API integrations
15+ hrs

manual work removed weekly

Python automation
25%

data-quality error reduction

programmatic validation
OPEN SOURCE & PERSONAL BUILDS

Projects I build when I want to explore.

Independent projects that show how I think outside day-to-day enterprise work — visualization, distributed systems, developer tooling and applied mathematics.

01 / DATA ENGINEERING
01

SQL Lineage Visualizer

Interactive table- and column-level SQL lineage with multi-dialect parsing, focus highlighting, query history, saved snippets and PNG export.

PythonFastAPISQLGlotReact Flow
02 / BIG DATA
02

SparkTune Pro

A high-performance Apache Spark cluster resource calculator designed to help Data Engineers reason about cluster sizing and configuration across YARN and Kubernetes.

Apache SparkReactVite
03 / WEB3 / STORAGE
03

Ether-Cloud

A decentralized cloud-storage application combining Ethereum smart-contract infrastructure with IPFS-based storage, developed around a blockchain-first architecture.

EthereumIPFSTruffleJavaScript
04 / AI / COMPUTER VISION
04

Image Fourier Transform

Streamlit application that converts uploaded images into frequency-domain representations, visualizes magnitude and phase spectra, and reconstructs images with the inverse Fourier transform.

PythonNumPySciPyStreamlit
04 TOOLBOX

Tools I Work With

The stack behind the systems I build — cloud, data, infrastructure and development tools, with their own visual language.

Hands-on across cloud + data
AWS
Terraform
PySpark
Apache Iceberg
Databricks
Python
SQL
Spark
Docker
React
Git
Azure
Hadoop
Scala
Airflow
LangChain
OpenAI APIs
RAG
Vector DBs
05 LET'S CONNECT

Let's Build Something.

I'm always open to interesting engineering opportunities, collaborations or simply a good conversation.