Abhay Kumar
Portrait of Abhay Kumar

LLM Research Engineer · Bengaluru, India

Abhay Kumar

I work on how large language models are trained: stability, optimisation and scale.

I worked on the Falcon-H1 models at the Technology Innovation Institute, and I am the first author of ZClip.

Falcon-H1
Worked on TII's hybrid Transformer and Mamba LLM family and the Falcon coder models.
ZClip
First author. Adaptive gradient clipping that stops loss spikes in LLM pre-training.
24B+ models
Hands-on LLM pre-training experience, up to models of 24B+ parameters.
GPT-2 in 2019
One of the first open-source GPT-2 implementations in TensorFlow 2.0. 267 GitHub stars, 78 forks.
01

Research

Google Scholar

  1. 2025arXiv preprintFirst author

    ZClip: Adaptive Spike Mitigation for LLM Pre-Training

    Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury, Fabian Güra

  2. 2025arXiv preprint

    A Refined Analysis of Massive Activations in LLMs

    Louis Owen, Nilabhra Roy Chowdhury, Abhay Kumar, Fabian Güra

  3. 2025arXiv preprint

    Variance Control via Weight Rescaling in LLM Pre-training

    Louis Owen, Abhay Kumar, Nilabhra Roy Chowdhury, Fabian Güra

  4. 2024arXiv preprint

    Komodo: A Linguistic Expedition into Indonesia's Regional Languages

    Louis Owen, Vishesh Tripathi, Abhay Kumar, Biddwan Ahmed

02

Open source

All repositories

  • TensorFlow 2★ 267

    gpt-2-tensorflow2.0

    One of the first GPT-2 implementations in TensorFlow 2.0, open-sourced in August 2019, before TensorFlow 2.0's stable release. Pre-training and text generation from scratch, with distributed training.

    akanyaani/gpt-2-tensorflow2.0

  • PyTorch★ 153

    ZClip

    Official implementation of the ZClip paper: adaptive gradient clipping from EMA statistics of the gradient norm, with optional PyTorch Lightning support.

    bluorion-com/ZClip

  • TensorFlow 2★ 41

    ranknet-tensorflow2.0

    Learning to rank, from RankNet to LambdaRank, implemented in TensorFlow 2.0.

    akanyaani/ranknet-tensorflow2.0

  • PyTorch★ 39

    miniLLAMA

    A compact implementation of the LLaMA and LLaMA 2 architectures for training and inference, written to make the differences from GPT easy to see.

    akanyaani/miniLLAMA

Star counts as of October 2026.

03

Experience

  1. Sep 2025 – May 2026

    Technology Innovation Institute

    Senior LLM Research Engineer · Abu Dhabi

    Worked on the Falcon-H1 models, TII's hybrid Transformer and Mamba LLM family, and on the Falcon coder models, across pre-training, evaluation and post-training refinement.

    Evaluation at scale. Designed and built a scalable evaluation pipeline for the coder models from scratch, on Kubernetes and Docker.

  2. Sep 2024 – Aug 2025

    BluOrion

    Senior LLM Research Engineer · Dubai

    Designed and led ZClip. Worked on initialisation and variance control, and built distributed pre-training pipelines on FSDP, PyTorch Lightning and DeepSpeed, used to train 1B to 13B parameter models.

  3. Dec 2020 – Sep 2024

    yellow.ai

    Research Scientist, NLP · Bengaluru

    Co-authored the Komodo LLM. Trained and deployed task-specific language models in production, and built zero-shot embedding and generative models that reduced unidentified utterances by 30%.

  4. Sep 2017 – Dec 2020

    EdGE Networks

    Senior Data Scientist, NLP · Bengaluru

    Early GPT pre-training. In 2019 I implemented GPT-2, a Transformer autoencoder and a custom autoregressive transformer from scratch in TensorFlow 2.0, and pre-trained them with distributed training on a large corpus.

    Open-sourced GPT-2 in 2019. I released it as gpt-2-tensorflow2.0, one of the first GPT-2 implementations in TensorFlow 2.0, published before TensorFlow 2.0's stable release. It now has 267 GitHub stars and 78 forks.

    Also built neural ranking and job recommendation models.

  5. May 2016 – Sep 2017

    Scry Analytics

    Data Scientist · Gurgaon

    Built opinion mining, named-entity recognition and phrase extraction models with LSTMs and CNNs, and a PySpark pipeline to run them on large document sets.

  6. May 2015 – May 2016

    Gauge Data Solutions

    Data Analyst · Noida

    Built a multithreaded web crawler for legal documents and a sequence classification model to clean the crawled data.

Education

  • 2009 – 2014Rajiv Gandhi Prodyogiki Vishwavidyalaya, Bachelor of Technology, Computer Science.
04

Contact

Working on LLM pre-training? Write to me at akanyaani@gmail.com.