Deep Learning Kernels — Accelerator Optimization

Jan 1, 2021 · 1 min read
projects

Project Details

01

Accelerator Framework

Developed a PyTorch and TensorFlow software suite for high-performance accelerator training and inference.
02

Kernel Debugging

Analyzed and resolved kernel bugs and performance bottlenecks.
03

TPC Mapping

Mapped PyTorch kernels to TPC kernels in C++.
04

Custom Operators

Developed operators using combinations of TPC kernels.
05

Model Export

Converted deep learning models to ONNX for deployment.
06

Graph Validation

Used Netron to visualize and validate exported model graphs.
07

Performance Focus

Optimized execution across accelerator training and inference workloads.

Tech Stack

Frameworks & Data

Python NumPy Pandas PyTorch TensorFlow

Acceleration & Tooling

C++ TPC Kernels ONNX Netron VS Code Linux Git JIRA
Shiv Kumar
Authors
Lead Engineer — GenAI / Agentic AI & Applied ML

Lead Engineer with 5 years of hands-on experience developing and deploying GenAI, RAG, Agentic AI, and deep learning systems from scratch. I build production-grade AI platforms for realtime voice interaction, enterprise knowledge retrieval, regulatory search, and automated finance analysis. My work spans NLP, autonomous driving, computer vision, and Camera-LiDAR-Radar fusion, with a focus on reliable architecture, model optimization, and scalable data pipelines.

Stack: Python, GCP, GenAI, LLMs, Agents, Docker, FastAPI, Cloud SQL, Google ADK, Vertex AI, Agent Engine, Cloud Run, OpenAI Realtime API, WebSockets, Redis, MongoDB, Milvus, JWT, Docling, BM25, ModernBERT, vLLM, and FAISS.