Available for opportunities

Govinda Neupane

I'm a

Building scalable modern data platforms using Snowflake, Spark, Hadoop, Python, and AWS.

Kathmandu, Nepal
Download Resume
Find me on
Govinda Neupane
GN
❄️

Snowflake

Certified

⚡

Big Data

Engineer

☁️

AWS EMR

Expert

5+
Years Exp.
7
Certs
15+
Projects
Scroll

About Me

Senior Data Engineer with expertise in building and migrating large-scale healthcare data platforms. Specialized in Snowflake, Snowpark Python, Hadoop, Spark, Java, and AWS EMR, with deep experience in distributed systems and HIPAA-compliant healthcare analytics. Currently leading a migration team modernizing legacy Hadoop pipelines to Snowflake through Augmented Engineering, driving code conversion, parity checking, and remediation of missing logic while ensuring data accuracy, performance, scalability, and reliability.

Engineering Philosophy

I believe in building data systems that are not just functional, but elegant — systems that scale gracefully, fail safely, and evolve with the business. Every pipeline I build is designed with observability, testability, and maintainability at its core.

Career Goals

Seeking senior-level opportunities in data engineering where I can architect and build next-generation data platforms, mentor engineering teams, and drive the adoption of modern cloud-native data technologies.

By the Numbers

0+
Years Experience
0+
Technologies Used
0+
Projects Delivered
0
Certifications

Areas of Expertise

Big Data Engineering

Hadoop, Spark, HDFS, MapReduce at petabyte scale

Cloud Data Platforms

Snowflake, AWS EMR, S3 — modern cloud-native pipelines

Pipeline Development

ETL, Snowpark Python, Java — end-to-end data flows

Healthcare Analytics

HCC risk adjustment, MARA, claims data processing

Data Quality

Validation, reconciliation, testing frameworks

Team Leadership

Code reviews, documentation, knowledge transfer

Education

🎓

BSc (Hons) in Information Technology

Asia Pacific University of Technology & Innovation (APU), Malaysia — studied at Lord Buddha Education Foundation, Kathmandu (APU-affiliated center)

Kathmandu, Nepal · 2018 - 2022

GPA: 3.01/4.0

Skills & Expertise

A comprehensive toolkit built through years of hands-on experience with enterprise-scale data systems.

🔍

Data Engineering

10 skills

Snowflake95%
Snowpark90%
Apache Spark88%
Hadoop85%
ETL Pipelines92%
Data Modeling85%
Data Validation90%
Data Reconciliation88%
Data Quality87%
Cascading75%

Programming

5 skills

Python92%
Java85%
SQL95%
Bash/Shell78%
TypeScript70%

Cloud & Infrastructure

5 skills

AWS EMR88%
AWS S385%
Elasticsearch80%
Docker75%
Git90%

Web Technologies

4 skills

Django82%
REST APIs85%
PostgreSQL83%
FastAPI72%

Soft Skills

6 skills

Collaboration95%
Documentation92%
Problem Solving95%
Code Reviews90%
Training & Mentorship88%
Agile/Scrum85%

Core Technology Stack

SnowflakeSnowparkApache SparkHadoopPythonJavaAWS EMRElasticsearchSQLDjangoPostgreSQLDockerGitREST APIs

Work Experience

5 companies · 8 roles · Building data platforms that scale

Leading the migration team modernizing large-scale healthcare data pipelines from legacy Hadoop systems to Snowflake, driving code conversion through Augmented Engineering while ensuring functional parity and data accuracy.

Key Achievements

  • Leading the migration team converting legacy big data pipelines to Snowflake using Augmented Engineering (AI-assisted) workflows
  • Driving code conversion from Hadoop (Java, Cascading, Spark) to Snowpark Python with accelerated, AI-augmented engineering practices
  • Performing rigorous parity checking between legacy and converted pipelines to guarantee functional equivalence and data accuracy
  • Identifying missing or incomplete business logic in AI-converted code and remediating gaps to ensure correctness
  • Reviewing converted Snowpark implementations and establishing engineering standards across the migration team
  • Coordinating cross-functional efforts to resolve data quality issues and keep the migration on track
  • Mentoring engineers and leading knowledge transfer on Snowflake, Snowpark, and Augmented Engineering workflows

Technologies

SnowflakeSnowparkPythonAugmented EngineeringApache SparkHadoopJavaCascadingSQLGit
5
Companies
8
Total Roles
3+
Years Active
Healthcare
Domain

Certifications & Achievements

Validated expertise through industry-recognized certifications from leading technology providers.

Featured Certifications
❄️
Featured

Snowflake Data Engineer

Snowflake

Advanced certification validating expertise in designing and building data engineering solutions on Snowflake.

2024
View
❄️
Featured

Snowflake Data Engineer II

Snowflake

Expert-level certification for advanced Snowflake data engineering patterns and architectures.

2024
View
🐍
Featured

Snowpark for Data Engineers

Snowflake

Specialized certification for building data pipelines using Snowpark Python and Java APIs.

2024
View
Additional Certifications
❄️

Snowflake Fundamentals

Snowflake

Foundational certification covering core Snowflake concepts, architecture, and features.

2023
View
🤖

Advanced Prompt Engineering

LinkedIn Learning

Advanced techniques for crafting effective prompts for large language models and AI systems.

2024
View
🐍

Learning Python

LinkedIn Learning

Comprehensive Python programming certification covering core language features and best practices.

2022
View
💻

Programming Foundations

LinkedIn Learning

Foundational programming concepts including algorithms, data structures, and software design principles.

2021
View
7 certifications earned and counting

Featured Projects

Real-world data engineering solutions built for scale, reliability, and performance.

Data EngineeringFeatured
In Progress

Hadoop to Snowflake Migration Platform

End-to-end migration framework for transitioning large-scale healthcare data pipelines from legacy Hadoop/MapReduce architecture to modern Snowflake cloud data platform. Includes automated validation, reconciliation, and rollback capabilities.

Migrated 50+ Hadoop MapReduce jobs to Snowflake
Achieved 99.9% data accuracy through automated reconciliation
Reduced pipeline execution time by 60%
SnowflakeSnowparkPythonApache SparkHadoop+2
Data QualityFeatured
Completed

Snowpark ETL Validation Framework

A robust testing and validation framework for Snowpark Python ETL pipelines. Provides automated data quality checks, schema validation, row count reconciliation, and statistical profiling for healthcare data pipelines.

100+ automated validation rules
Reduced data quality issues by 80%
Integrated with CI/CD pipelines
PythonSnowparkSnowflakepytestGreat Expectations+1
Healthcare AnalyticsFeatured
Completed

Healthcare Claims Analytics Engine

High-performance analytics engine for processing and analyzing healthcare claims data at scale. Supports HCC risk adjustment calculations, MARA analytics, and utilization metrics for value-based care programs.

Processes 10M+ claims records daily
Sub-second query response with Elasticsearch
Supports HCC and MARA risk models
JavaApache SparkHadoopAWS EMRElasticsearch+2
Data Engineering
Completed

Spark-based Data Processing System

A scalable, fault-tolerant data processing system built on Apache Spark for batch and streaming data workloads. Features dynamic resource allocation, job monitoring, and automated retry mechanisms.

Handles 1TB+ daily data volumes
99.5% pipeline reliability
Automated failure recovery
Apache SparkPythonScalaAWS EMRApache Airflow+1
+

More projects coming soon

Edit src/data/projects.json

Let's Connect

Open to senior data engineering roles, consulting opportunities, and interesting data platform challenges.

Contact Information

Email

govindaneupane81@gmail.com

Location

Kathmandu, Nepal

Available for Hire

Open to senior data engineering roles and consulting. Response within 24 hours.

Send a Message