Open to new opportunities

Omar Waseem

AI Engineer  ·  Data Engineering Foundation  ·  Eagle Scout

I'm an AI Engineer with a data engineering foundation. I started out at Revature building distributed ETL pipelines on Spark and Hadoop, then moved into an AI Engineering role at Cognizant designing multi-agent LLM workflows and forecasting models. I like grounding AI systems in solid data fundamentals: retrieval-augmented reasoning, agent workflows with real guardrails and human-in-the-loop gates, and pipelines that hold up under pressure. Outside of work, I'm an Eagle Scout and still carry that build-it-right mentality into everything I do. Feel free to check out my projects, browse my experience, or grab my resume below.

Experience

Cognizant
AI Engineer Intern
May 2026 – Aug 2026
  • Built a Python ETL pipeline extracting flight, hotel, and live weather data, validating and cleaning it, and loading it into PostgreSQL staging tables, paired with a Flask dashboard for pipeline health and reporting metrics.
  • Built a pooled LSTM network (TensorFlow/Keras) with learned per-stock embeddings to forecast next-day returns across ~100 large-cap equities, benchmarking against baseline models and validating results with a capacity-ablation study.
  • Designed a multi-agent LangGraph/LangChain workflow automating vendor onboarding and risk review — document validation, RAG-grounded policy retrieval, risk scoring, and a human-approval gate — with a Streamlit interface and full scenario test coverage.
Revature
Data Engineer Intern
Feb 2026 – May 2026
  • Completed a 9-week intensive data engineering track (SQL, Git, Python, Hadoop, Hive, PySpark) spanning 58 coding labs, including distributed ingestion and transformation workflows across HDFS, Hive, and Spark SQL on AWS EMR clusters.
  • Built a Python data pipeline as a capstone project, ingesting and validating CSV data into a SQLite database and writing SQL queries to aggregate analytics.
Xtremeanalytix
Cloud Engineering Intern
Jun 2024 – Aug 2024
  • Designed and deployed five data pipelines across SQL Server, MongoDB, AWS RDS, and Azure SQL in cloud and hybrid environments, accelerating data integration for downstream applications.
  • Created a custom MongoDB source connector in Java enabling real-time data streaming into Apache Kafka, reducing data ingestion latency by 40% across 3 cloud-hosted applications.
  • Optimized data synchronization across 4+ cloud-native and on-prem systems, improving data accessibility and reducing pipeline errors by 30%.
  • Maintained cloud infrastructure on AWS and Azure, supporting fault-tolerant ingestion workflows handling millions of records daily.

Education

Rutgers University
Bachelor of Science, Computer Science
New Brunswick, NJ  ·  May 2025

Projects

🛂
LangGraph + LangChain multi-agent workflow that runs a vendor request through document/completeness validation, RAG-grounded risk & compliance review, and a human-approval gate before producing a grounded onboarding decision. Model-agnostic LLM client (Bedrock → OpenAI → Anthropic → Gemini) with deterministic template fallback, a Chroma vector store, and a Streamlit demo UI.
3-agent workflow 5 required test scenarios
LangGraph LangChain RAG Python
🧠
Pooled LSTM network (TensorFlow/Keras) with learned per-stock embeddings predicting next-day returns across ~100 large-cap US equities from live Yahoo Finance data. Benchmarked against baseline models with an honest read on real-world predictive value, backed by a capacity-ablation study confirming the performance ceiling is data-driven, not architecture-driven. Ships with a Flask web app for live predictions.
~100 stocks modeled Capacity-ablation validated
TensorFlow / Keras Python LSTM Flask
🛫
Transportation data pipeline that extracts flight, hotel, and live weather data across 98 cities, validates and cleans it against config-driven rules, and loads it into PostgreSQL staging tables with FK-based reject handling. Paired with a Flask + Chart.js dashboard for pipeline stats, a city-level reporting view, and a searchable rejects table.
87% test coverage 98 cities tracked
Python PostgreSQL Flask pandas
☁️
Production-style serverless pipeline using AWS SAM, API Gateway, and Lambda. Integrates Claude Haiku 4.5 via Amazon Bedrock for automated text summarization and sentiment tagging, with results persisted in DynamoDB and all endpoints secured via API key authentication.
61% text reduction 1.89s avg response 1,000+ requests
AWS Lambda DynamoDB Bedrock / Claude Python
📈
Production-style RESTful API in ASP.NET Core (.NET 10) with 16 JWT-secured endpoints across Auth, Portfolio, Logs, and Admin controllers. Integrates Alpha Vantage for live stock pricing with 60-second in-memory caching. Designed for Azure App Service deployment with EF Core migration swap.
16 JWT endpoints 80% fewer API calls
ASP.NET Core C# Azure SQLite
⚙️
Open-source Go CLI tool for declarative cross-platform environment management across macOS, Linux, and WSL2. Contributed the macOS and WSL2 side of development, including a Homebrew service wrapper and WSL2 PATH sanitization logic to prevent Windows binary shadowing. Full CI/CD via GitHub Actions across 13 packages.
13 packages Full integration tests
Go CLI GitHub Actions Open Source
🔄
Java-based Kafka connector enabling real-time streaming from MongoDB to Apache Kafka. Built for hybrid cloud integration with AWS and Azure, cutting data ingestion latency by 40% across production cloud-hosted applications.
40% latency reduction
Java Apache Kafka MongoDB AWS / Azure
🎮
Interactive Pokémon-themed Wordle bot supporting daily and personal game states for 100+ active users. JSON-based persistence with optimized command handling. Deployed to Render with UptimeRobot for automated uptime monitoring.
100+ active users <200ms response
Python Discord API Render
💊
Cross-platform mobile app with JWT authentication, SQLite storage, and dark mode. AI recommendation engine using Bayesian reasoning and informed search, delivering personalized supplement suggestions with confidence scores.
250ms response time
React Native Express.js Bayesian AI SQLite
🧭
AI agent using reinforcement learning and Q-learning to navigate a custom maze environment. Implements fast replanning with a custom reward function to efficiently recover from obstacles.
Python Reinforcement Learning Q-Learning
🧬
Python simulation using NumPy and Matplotlib to visualize cellular automata. Demonstrates algorithm design and efficient matrix operations for cellular state transitions.
Python NumPy Matplotlib
🗃️
Custom malloc/free memory allocator in C with multiple test cases demonstrating systems-level understanding of heap allocation strategies and memory management.
C Systems Memory Management

Technical Skills

Languages
Python Go C# Java JavaScript C / C++ HTML / CSS
AI / ML
LangGraph LangChain RAG TensorFlow / Keras scikit-learn Chroma / Vector DBs Bedrock
Cloud & Data
AWS Lambda API Gateway DynamoDB RDS CloudWatch Azure Docker Apache Kafka Apache Hadoop PySpark MongoDB PostgreSQL SQL / SQLite pandas ETL
Frameworks & Tools
React React Native Express.js ASP.NET Core Entity Framework Flask Streamlit AWS SAM GitHub Actions

Certifications

🏅
AWS Certified AI Practitioner
Amazon Web Services · 2026
🎓
Fundamentals of AI Agents Using RAG and LangChain
IBM · Coursera
🎓
Machine Learning Specialization
DeepLearning.AI · Coursera
⚜️
Eagle Scout
Boy Scouts of America · Earned by fewer than 6% of Scouts