AI engineer building LLM systems, RAG pipelines, and 3D vision tools.

MSc Artificial Intelligence, University of Surrey. I focus on backend systems, evaluation, and deployment.

Open to AI roles and freelance work

Selected work

All projects (8)

Log Guardian

Built

Evidence-based security log review with optional AI

  • Python
  • FastAPI
  • PostgreSQL
  • SQLite
  • OpenAI API
  • Playwright
  • Docker

Built a review tool that links gateway requests to explicit authentication outcomes and produces cited facts without a model. Optional AI selects bounded evidence and allowed explanations; invalid selections are rejected without changing the factual baseline. Tested five browser-driven scenarios against a locally instrumented OWASP Juice Shop, checking 182 records across reviews representing 154 unique source records. Missing authentication logs stayed unknown despite HTTP success. Includes immutable case snapshots, separate review/execution permissions, and preserved failure evidence. Alpha prototype, not a validated attack detector or production security service.

View project: Log Guardian

Mini OpenAI Platform

Built

Self-hosted LLM platform with OpenAI-compatible APIs

  • Python
  • FastAPI
  • React
  • Qdrant
  • Ollama
  • Docker Compose
  • Prometheus
  • Grafana

Microservices LLM platform — API gateway, RAG, embedding, and inference services — with an embedding-based semantic cache cutting latency ~80× on cache hits, a difficulty-based router dispatching prompts across Ollama model tiers, a CI quality gate on retrieval metrics (recall@k, MRR), and 16 Prometheus/Grafana panels covering latency, token economics, and answer quality.

View project: Mini OpenAI Platform

Chat2Study

Built

Turn AI chat transcripts into a searchable knowledge base

  • Next.js
  • TypeScript
  • FastAPI
  • LangGraph
  • PostgreSQL
  • pgvector
  • MinIO
  • Docker

Full-stack RAG app converting long AI chat transcripts into knowledge bases, study notes, and concept maps. A 10-node LangGraph pipeline orchestrates Playwright capture, artifact persistence, chunking, and embedding; async job re-architecture took ingestion API responses from 30–90s to ~10ms, with pgvector retrieval at ~50ms, a provider-agnostic LLM factory, JWT auth, and full CI.

View project: Chat2Study

Selected writing

All writing

About

I'm Hitendra, an AI engineer based in the United Kingdom. My work spans LLM infrastructure, retrieval systems, and interactive 3D vision. For my MSc at the University of Surrey, I built scribble-guided segmentation tools for 3D Gaussian scenes.

I also enjoy learning languages. I work in English and Spanish, speak conversational Mandarin, and grew up speaking Hindi and Marathi.

English
Fluent
Spanish
Fluent
Mandarin
Conversational
Hindi
Native
Marathi
Native
Japanese
Basic
German
Basic
Norwegian
Basic

Alongside full-time AI roles, I'm open to scoped freelance work and research collaborations. Get in touch.