Open to UK data engineering roles

Pranava Kailash Subramaniam Prema

Data engineer.Reliable pipelines.Useful data.

I build Python and SQL ETL pipelines for regulatory data, schedule data workflows with Apache Airflow, and work with AWS ingestion and storage.

01 / Profile

From source data to something useful.

I build data pipelines that turn messy inputs into datasets people can use. My background includes commercial ETL work and an MSc in Data Science from the University of Surrey.

At Fleet Street Research, I worked on Python and SQL pipelines for regulatory data: collecting records, transforming them across MySQL and SQLite, and checking for missing data and schema changes. I also built Power BI semantic models and DAX measures for risk and compliance reporting.

My independent projects use Apache Airflow for scheduled data workflows and AWS Kinesis, EMR, and S3 for weather-data ingestion and processing. Machine learning remains part of my background, but data engineering is where I want to focus.

01

Make source data usable

Ingest records, clean them, and keep transformations traceable for the people who use the results.

02

Check before publishing

Catch schema changes, missing records, and processing errors before they affect reporting.

03

Make runs repeatable

Use schedules, retries, and clear documentation so a pipeline can run beyond a single demo.

02 / Experience

Work first. Education in context.

Oct 2022 — Jan 2024

Fleet Street Research · Freelance · UK

Regulatory Data Researcher / Engineer

  • Maintained Python and SQL ETL pipelines across MySQL and SQLite for compliance datasets.
  • Built validation checks for schema changes, missing records, and processing exceptions in AML, CFT, and KYC data.
  • Improved data processing time by approximately 50% and built Power BI semantic models and DAX measures for compliance reporting.

Feb 2022 — Apr 2022

Yoshops.com · India

Data Scientist Intern

  • Prepared sales and pricing data with Python and SQL; pricing models improved accuracy by approximately 40%.

Mar 2021 — Jun 2021

Forsk Coding School · India

Data Scientist Intern

  • Integrated sentiment analysis outputs with a Flask backend and deployed the solution on AWS.
03 / Selected work

Data pipelines first.

I put the data engineering work first: ingestion, orchestration, storage, and validation. The ML and product projects show where that work has taken me next.

01

Data engineering · Orchestration

Financial Data Pipeline

A scheduled workflow that collects NVIDIA price history and financial news for a dashboard of indicators and sentiment.

What I built
Built a weekday Apache Airflow DAG with validation and retries, and packaged the dashboard, Airflow, and its metadata database with Docker Compose.
Evidence
The repository documents the workflow, weekday schedule, and integrity checks.
  • Python
  • Apache Airflow
  • pandas
  • Docker
  • Dash
  • PostgreSQL

02

Data engineering · ETL

Weather Data ETL

A containerised pipeline that extracts weather observations from OpenWeather, transforms the API response, and loads it into PostgreSQL.

What I built
Wrote the Python ETL task as an Apache Airflow DAG in an Astro project, with Docker for repeatable local setup.
Evidence
The public repository includes the DAG, setup steps, and Airflow and database screenshots.
  • Python
  • Apache Airflow
  • Astro
  • PostgreSQL
  • Docker

03

Data engineering · AWS project

Weather Sensor Pipeline

An independent project connecting streaming weather-sensor ingestion, batch processing, and cloud storage.

What I built
Sent sensor data through Amazon Kinesis, processed batches with Amazon EMR, and stored the results in Amazon S3.
Evidence
Independent project; implementation details are available on request.
  • Python
  • Amazon Kinesis
  • Amazon EMR
  • Amazon S3

04

Machine learning · MSc dissertation

CyNER 2.0

A cybersecurity named-entity recognition system for finding malware, vulnerabilities, threat actors, systems, and organisations in text.

What I built
Compared five model approaches and exposed the selected model through a FastAPI service.
Evidence
91.88% F1 score on augmented data; more than 11,000 annotated rows.
  • Python
  • DeBERTa
  • FastAPI
  • SQLAlchemy
  • SQLite

05

Independent product · On-device AI

Prompt Enhancer

A Chrome extension that rewrites prompts inside ChatGPT, Gemini, Claude, and Perplexity using Chrome’s on-device Gemini Nano model.

What I built
Built the browser integration, two enhancement modes, local history, keyboard shortcuts, and limited host permissions.
Evidence
Prompts stay on-device; the extension does not send them to a server.
  • Chrome API
  • Gemini Nano
  • Manifest V3
04 / Capabilities

Tools connected to work.

These are tools I have used in client work or independent projects. Each area points back to a pipeline, report, or system on this page.

01

ETL and orchestration

Python · SQL · Apache Airflow · MySQL · SQLite

Commercial compliance-data pipelines and scheduled financial and weather workflows.

02

Cloud data

AWS Kinesis · Amazon EMR · Amazon S3 · Docker

Independent weather-sensor pipeline from streaming ingestion to batch processing and storage.

03

Analytics and quality

Power BI · DAX · pandas · Data validation

Fleet Street compliance reporting, semantic models, and checks for schema drift and missing records.

04

Supporting ML work

scikit-learn · Hugging Face Transformers · FastAPI

Cybersecurity entity recognition, evaluation, and API-based inference.

06 / Contact

Building a data engineering team?

I am based in Guildford and open to UK data engineering roles with Skilled Worker sponsorship.

Guildford, United Kingdom

pranavakailashsp@gmail.com

Available for full-time roles