You're seeing this page as if you were . The main menu is still yours, though. Exit from immersion
Vishal AwateVA

Vishal Awate

LEAD DATA ENGINEER · DATA PLATFORMS · GenAI / LLM

€600/day
Pune, IN
8-15 years

Average response time: 1 hour

About Vishal

Lead Data Engineer with 9+ years of experience designing petabyte-scale data platforms — and a hands-on builder of production GenAI/LLM systems.

I build GenAI for a global banking client — a production service that translates complex SQL and multi-language code (Python, Scala, Java, and more) into plain-English business explanations: hybrid RAG combining deterministic structured retrieval with vector search, database-driven prompt management, LLM-as-judge quality evaluation, and fully deterministic outputs built for regulated environments, running on Llama 3.3 through an enterprise LLM gateway.

Independently, I designed and operate a second production GenAI system end to end — multi-source data ingestion (25+ sources), an LLM pipeline with structured extraction using tool schemas, embedding-based semantic search (RAG) with vector storage in PostgreSQL, multi-stage match scoring with model cascading that keeps inference costs under control, LLM-as-judge evaluation gating notifications, and LLM-generated documents delivered over WhatsApp and email.

I architect end-to-end data solutions: petabyte-scale data lake pipelines in the payments domain (PayPal/Braintree) using Spark/PySpark, lakehouse platforms on Azure Databricks, and the orchestration, quality, and cost controls that keep them reliable in production — plus platform governance frameworks automating storage, compliance, and cost monitoring across 10+ Hadoop databases serving 15+ teams.


Tech: Spark · PySpark · Python · Azure Data Factory · Azure Databricks · AWS (S3, Redshift, Lambda) · GCP core services · Hive · Impala · HDFS · SQL (MySQL, Oracle, PostgreSQL) · Shell · CI/CD (Jenkins, uDeploy)

AI: LLMs (Claude, OpenAI, Llama 3.3) · Hybrid RAG (structured + vector) · Embeddings & Vector Search (pgvector) · AI Agents · Prompt Engineering · LLM-as-Judge · Model Cost Optimization · Deterministic LLM Outputs
  • English

    Native or bilingual

Remote only
Primarily works remotely

Experience

  • Atyeti Inc
    Lead Associate — Data Engineering Lead
    May 2024 - Today (2 years and 3 months)
    Pune, Maharashtra, India
    Lead Data Engineer with 9+ years of experience designing petabyte-scale data platforms — and a hands-on builder of production GenAI/LLM systems.

    I build GenAI for a global banking client — a production service that translates complex SQL and multi-language code (Python, Scala, Java, and more) into plain-English business explanations: hybrid RAG combining deterministic structured retrieval with vector search, database-driven prompt management, LLM-as-judge quality evaluation, and fully deterministic outputs built for regulated environments, running on Llama 3.3 through an enterprise LLM gateway.

    Independently, I designed and operate a second production GenAI system end to end — multi-source data ingestion (25+ sources), an LLM pipeline with structured extraction using tool schemas, embedding-based semantic search (RAG) with vector storage in PostgreSQL, multi-stage match scoring with model cascading that keeps inference costs under control, LLM-as-judge evaluation gating notifications, and LLM-generated documents delivered over WhatsApp and email.

    I architect end-to-end data solutions: petabyte-scale data lake pipelines in the payments domain (PayPal/Braintree) using Spark/PySpark, lakehouse platforms on Azure Databricks, and the orchestration, quality, and cost controls that keep them reliable in production — plus platform governance frameworks automating storage, compliance, and cost monitoring across 10+ Hadoop databases serving 15+ teams.


    Tech: Spark · PySpark · Python · Azure Data Factory · Azure Databricks · AWS (S3, Redshift, Lambda) · GCP core services · Hive · Impala · HDFS · SQL (MySQL, Oracle, PostgreSQL) · Shell · CI/CD (Jenkins, uDeploy)

    AI: LLMs (Claude, OpenAI, Llama 3.3) · Hybrid RAG (structured + vector) · Embeddings & Vector Search (pgvector) · AI Agents · Prompt Engineering · LLM-as-Judge · Model Cost Optimization · Deterministic LLM Outputs
    LLM Apache Spark PostgreSQL RAG AI Agent
  • Clairvoyant LLC (an EXL Company)
    Senior Software Engineer
    December 2017 - April 2024 (6 years and 4 months)
    Pune, Maharashtra, India
    • • Core engineer on PayPal/Braintree's enterprise data lake (CPDL) — petabyte-scale payments data across ingestion, processing, and data-quality layers (Braintree, Platform, and Swift workstreams).
    • • Built a generic PySpark batch-ingestion framework with integrated data-quality checks (null, duplicate, trend analysis), adopted as the team standard for onboarding new hourly data sources — cutting source onboarding from custom builds to configuration.
    • • Developed AWS extraction pipelines (S3, EC2, Redshift, Boto) feeding the CPDL batch-ingestion framework, with UC4-scheduled daily production workflows — decryption and validation before ingestion into Hive, with transformations on top.
    • • Designed a CDM automation tool that mapped multiple source tables into Common Data Model tables automatically, eliminating repetitive manual modelling work.
    • • Optimized long-running Spark/Hive workloads via partitioning, bucketing, and file-size management on CPDL's multi-terabyte daily volumes.
    • • Delivered Lloyds Bank's mainframe-to-Hadoop migration, moving legacy banking data into HDFS/Hive with reconciliation checks.
    • • Led Azure cloud migration for a cybersecurity client (Abacus) — metadata-driven Azure Data Factory pipelines ingesting REST APIs, PostgreSQL, MySQL, and SQL Server into a medallion architecture (bronze/silver/gold) with Databricks notebook transformations, Key Vault-managed credentials, gold-layer data models, and Grafana monitoring.
  • Centurysoft Private Limited
    Hadoop Developer
    May 2017 - October 2017 (5 months)
    Pune, Maharashtra, India
    Social-Media Analytics Client
    • • First industry role in big data — built Hadoop-ecosystem data-processing jobs (HDFS, Hive, Sqoop, Pig), processing terabytes of daily log data for a social-media analytics client (Solvedrum) into controls and change-history reporting.

Recommendations

Be the first to recommend Vishal

Help this freelancer shine by sharing your experience working together.

These freelancer profiles also match your criteria

AgathaA

Agatha Frydrych

Backend Java Software Engineer

4.7

(3)

2

BaptisteB

Baptiste Duhen

Fullstack developer

4.6

(4)

5

AmedA

Amed Hamou

Senior Lead Developer

4

(2)

7

AudreyA

Audrey Champion

Web developer

4.3

(3)

4

Education

  • B.Tech
    Shivaji University
    B.Tech

Skill set

Categories