Srikar Palani

software engineer

I'm studying at the University of Illinois Urbana-Champaign and building systems across machine learning, infrastructure, and reliability. Recently, I built a production ML inference reliability platform at Meta and infrastructure-security systems at AWS.

Experience

Meta

Software Engineering Intern

ML Inference Platform Efficiency · Menlo Park, CA · June–August 2026

Built a production ML inference reliability platform across tens of thousands of serving jobs, reducing investigations from days or weeks to approximately 30 minutes. Architected a deterministic execution engine for AI-generated studies with schema validation, provenance caching, race-safe concurrency, and asynchronous scheduling. Developed LLM-agent workflows with 91% precision and 100% recall on routing evaluation, productionizing 876 runs and approximately 1.47 million classifications in two weeks.

Amazon Web Services

Software Development Intern

Infrastructure Security · Seattle, WA · May–August 2025

Built and deployed a production infrastructure-security platform managing millions of network devices across Amazon's global infrastructure. Designed serverless Python APIs with AWS Lambda, API Gateway, and Neptune to automate ACL validation, access-path verification, and role-based authorization.

Infocrunch

Software Development Intern

Remote · June–August 2024

Built production Flutter and React applications and a FastAPI/PostgreSQL data pipeline supporting workforce management across 3,500+ rural banks, including attendance, transaction, account, and payout tracking.

Projects

GPT-2 Language Model from Scratch (PyTorch)

Implementing a decoder-only transformer from first principles, including causal multi-head self-attention and autoregressive generation. Building a training pipeline with tokenization, AdamW, learning-rate scheduling, gradient accumulation, mixed precision, and checkpointing. Developing and benchmarking KV caching and batched decoding to study latency, throughput, and memory tradeoffs.

High-Performance ML Kernels (CUDA, Tensor Cores)

Implemented CUDA WMMA convolution kernels using TF32 Tensor Cores on NVIDIA A40 GPUs. Fused convolution operations and optimized global-memory access, shared-memory tiling, warp execution, and thread-block decomposition to improve GPU throughput.

UNIX-like RISC-V Operating System

Built a Unix-like RISC-V operating system with virtual memory, demand paging, system calls, ELF loading, and user/kernel isolation. Developed a block-based filesystem with caching and VIRTIO-backed storage, plus preemptive scheduling, context switching, pipes, and inter-process communication.

Education

University of Illinois Urbana-Champaign

Champaign, IL · May 2027

M.S. in Computer Science
B.S. in Computer Engineering · GPA: 3.81/4.0

Coursework: Distributed Systems, High Performance Parallel Computing, Applied Machine Learning, Compilers, Parallel Programming, Operating Systems, Networking,

Honors: James Scholar, Dean's List

Skills

Languages
Python, C++, C, Java
ML / Systems
PyTorch, CUDA, Tensor Cores, Pandas, OpenCV, Linux, Parallel Programming, Git
Infrastructure
AWS (Lambda, API Gateway, Neptune, CDK, CloudFormation, S3, VPC, IAM), FastAPI, PostgreSQL