Revanth Posina
Data Engineer (AI/ML) at Microsoft.
I build the pipelines, forecasts and LLM agents that power finance analytics at Microsoft.
About
I turn messy operational data into numbers people plan on, and I build the agents that explain them.
Since 2020 I've moved from gaming telemetry at Entain to healthcare claims at Bloom Insurance to finance analytics at Microsoft. The domains change, the work doesn't: get the data in reliably, model it so the numbers hold up, and make it easy for people to ask why.
I care most about the parts that make data trustworthy: clear contracts between teams, quality checks, lineage, and agents with guardrails and evals. Forecasts and anomaly detection are only useful when people believe the data underneath.
I studied data science at Indiana University Bloomington. Away from the keyboard: trail running, hiking and Formula 1.
Work and projects
A distributed analytical query engine built from first principles, to show the planning, shuffles, scheduling and fault tolerance that Spark, Trino and DuckDB normally hide.
What it does
- A coordinator parses SQL into logical and physical plans, applies optimizer rules and schedules a distributed execution DAG across worker nodes.
- Operators for scans, filters, projections, aggregations, sorts and hash joins. Broadcast or shuffle hash join, chosen from table statistics and configurable thresholds.
- A custom shuffle layer with hash and range partitioning, bounded buffers, backpressure and spill to disk when memory runs short.
- Its own catalog of partitions, row counts and min/max stats, for partition pruning, predicate pushdown and projection pruning.
- Heartbeats and task state on the coordinator. When a worker dies, its unfinished partitions are reassigned and retried. Failure-injection tests kill workers mid-scan, mid-shuffle and mid-aggregation and check the results stay correct.
What I'm measuring
Not built to compete with production databases. It's here to show the systems fundamentals underneath them. Benchmarks publish with the repo.
A streaming lab that detects, diagnoses and recovers from data and infrastructure failures, benchmarked on a ladder from 100M up to 1.2B events an hour.
What it does
- Load generators drive a multi-broker Kafka cluster with configurable event sizes, key skew, bursts and schema versions.
- Spark Structured Streaming or Flink validates, enriches and aggregates into Iceberg or Delta. Bad events go to a replayable dead-letter queue.
- A reliability control plane runs detect, diagnose, decide, remediate, validate, escalate. Safe actions go through a bounded policy engine; risky ones, like breaking schema changes, get quarantined and escalated.
- A chaos framework injects broker loss, hot partitions, skew, memory pressure, slow sinks, and late or duplicate events, then records time to detect and recover, loss and duplicate rates.
- An optional AI reliability agent writes root-cause hypotheses but cannot run infrastructure commands. Every action passes through deterministic policy.
What I'm measuring
These are targets, not results yet. Numbers go up when the benchmarks run.
The data systems behind model development: reproducible, versioned training datasets and an evaluation platform with regression gates. Working name TrainForge.
What it does
- Ingests documents, app events, structured records, model outputs and human feedback into an immutable raw lakehouse layer.
- Distributed preprocessing: schema validation, normalization, language detection, PII filtering, quality scoring, tokenization, and semantic dedup with MinHash and LSH.
- A Dataset Registry versions every dataset with its source snapshots, transforms, filters, tokenizer and schema versions, and lineage, so any experiment can be reproduced exactly.
- Point-in-time snapshots, incremental processing, and safe backfills when a quality rule changes.
- Eval workers score models, prompts and retrieval strategies on accuracy, similarity, assertions, LLM-as-judge, latency, tokens and cost. Regression gates block promotion below thresholds, and a contamination check catches train and eval overlap.
What it enforces
In design and early build. Drift monitoring and freshness metrics land in Prometheus and Grafana.
Natural-language analytics over a governed warehouse: a text-to-SQL agent grounded in a semantic layer of metric definitions and schema metadata.
What it does
- Grounded in a semantic layer of metric definitions and schema metadata rather than raw tables.
- Generated SQL is parsed and validated before execution, runs under read-only credentials with row limits, and self-corrects on errors.
- Evaluated on a gold question set against a schema-only baseline: execution accuracy, p95 latency and cost per query.
Eval
Built on public TPC-H data and separate from my Microsoft work. Eval results go up here once the run finishes.
Explainable risk scoring on about 430K CDC BRFSS 2023 survey responses, deployed as a low-latency SageMaker endpoint.
What it does
- Mapped 350+ survey fields through the BRFSS codebook down to 33 curated features, with log and z-score transforms on skewed fields and statistical ranking using Cohen's d, t-tests and chi-squared.
- Compared XGBoost against LightGBM and logistic regression, tuned it with 150+ Optuna trials, and used SHAP for global and per-person explanations.
- Checked robustness by retraining without the dominant prior-diagnosis feature.
- Deployed on a SageMaker endpoint at under 150 ms p95, with CloudWatch latency alarms, MLflow tracking and an Airflow DAG from data prep through deploy.
- A Streamlit app for CSV upload, live scoring and SHAP views.
Results
A serverless AWS pipeline that pulls Spotify track, artist and album data every day and makes it queryable in Athena.
What it does
- A Lambda function pulls from the Spotify API on a daily CloudWatch Events schedule and lands raw data in S3.
- S3 events trigger a transform Lambda that cleans the data, checks the schema and writes to a transformed zone.
- Glue Data Catalog for schemas and Athena for SQL. Snowflake via Snowpipe is the planned next step.
Results
An n8n agent that checks my calendar, the weather and air quality, then emails me the best trail for the day, or a reason to skip it.
What it does
- Five tools: checkCalendar, getWeather, getAirQuality (AirNow PM2.5), getHikeList from a trail sheet, and sendEmail.
- Recommends a trail only when the weather and air quality are good, and picks one that fits my time, elevation and shade preferences.
- Otherwise it sends a heads-up with the reason, so every decision is explainable.
Results
A 3D open-world platformer scene inspired by the classics, built in Unity and C#.
What it does
- A 3D scene that takes the classic platformer idea to an open-world scale.
- Player interactions and character animations, with TextMeshPro for in-game text.
- A gameplay video lives in the repo README. The goal is a fully interactive open world with multiple levels.
Experience
Microsoftcurrent
Building the data foundation, forecasting and LLM agents behind finance analytics.
Project990
Turned IRS Form 990 filings into clean datasets and a search service for analysts.
Bloom Insurance
Migrated claims analytics to a streaming cloud platform and kept its data trustworthy.
- Data Engineer IJul 2024 – Dec 2024
- Data Engineer InternMay 2023 – Dec 2023
Ivy Comptech (Entain)
Built real-time and batch data pipelines for gaming analytics and set data contracts across producer teams.
- Data Engineer, OpsFeb 2021 – Jul 2022
- Trainee Software Engineer, Data OpsJun 2020 – Feb 2021
Education and more
Stack
Languages
7What I write every day, plus what I use for systems and game work.
Processing and streaming
5Batch and streaming engines.
Lakehouse and storage
6Open table formats, columnar files and object storage.
Warehousing and modeling
7Where curated data lives, and how it is shaped.
Orchestration and quality
6Scheduling, data contracts, tests and CI/CD.
ML and MLOps
8Models from features to monitored endpoints.
GenAI and agents
8Retrieval and agents, with guardrails and evals.
Infra and observability
7Running it, and knowing when it breaks.
Off the clock
Also into
Working on forecasting, streaming or LLM systems? Let's talk.
Email is the fastest way to reach me.