LLM pre-training · Bengaluru, India
Abhay Kumar
I work on how large language models are trained: stability, optimisation and scale.
Co-founder, FrontiersMind, building open foundation models that are sovereign and efficient. Author of ZClip.
About
I am an LLM researcher and the co-founder of FrontiersMind. My work is on pre-training: what makes large runs diverge, how to initialise and control variance so they do not, and how to build the distributed pipelines that train them.
Before FrontiersMind I was a Senior LLM Research Engineer at the Technology Innovation Institute in Abu Dhabi, working on the Falcon LLMs, and at BluOrion in Dubai, where I designed ZClip, an adaptive gradient clipping algorithm that removes the need to skip batches by hand, and worked on initialisation and variance control. I have co-led the pre-training of 7B and 13B parameter models on 15 trillion tokens. Earlier, at yellow.ai, I co-authored the Komodo LLM for Indonesia's regional languages.
I have worked in machine learning and NLP for more than ten years. Several from-scratch implementations I wrote along the way, of GPT-2 in TensorFlow and of LLaMA in PyTorch, are open source.
- 7B · 13B
- parameter models whose pre-training I co-led
- 15T
- tokens in that pre-training run
- FSDP · DeepSpeed · Lightning
- distributed training stacks I build pipelines on
-
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
An adaptive gradient clipping algorithm that sets its threshold from running statistics of the gradient norm, using z-score anomaly detection to catch the spikes that cause loss blow-ups without interfering with convergence otherwise.
-
A Refined Analysis of Massive Activations in LLMs
An analysis of massive activations across a broad range of LLM architectures. Not all of them are detrimental, and known mitigations are model-specific; pairing Target Variance Rescaling with Attention KV bias or Dynamic Tanh balances mitigation with downstream performance.
-
Variance Control via Weight Rescaling in LLM Pre-training
Introduces Layer Index Rescaling for weight initialisation and Target Variance Rescaling for variance control. On a 1B parameter LLaMA model, better variance management improves downstream performance by up to 4.6% and reduces extreme activation values.
-
Komodo: A Linguistic Expedition into Indonesia's Regional Languages
Komodo-7B, a family of 7-billion-parameter LLMs that work across Indonesian, English and 11 regional languages of Indonesia.
-
BED: Bi-Encoder-Based Detectors for Out-of-Distribution Detection
Bi-encoder-based detectors for out-of-distribution detection in NLP, with a comparison of OOD methods and feature extractors on CLINC150, ROSTD-Coarse, SNIPS and YELLOW.
-
PythonNew, 2026
faithserve
Is this endpoint serving this open-weight model faithfully? A CLI that runs black-box checks against an OpenAI-compatible server: chat template parity, sampling parameters, streaming and tool calls.
FrontiersMindAI/faithserve
-
PyTorch★ 153
ZClip
Official implementation of the ZClip paper: adaptive gradient clipping from EMA statistics of the gradient norm, with optional PyTorch Lightning support.
bluorion-com/ZClip
-
TensorFlow 2★ 267
gpt-2-tensorflow2.0
OpenAI GPT-2 pre-training and sequence prediction, implemented in TensorFlow 2.0.
akanyaani/gpt-2-tensorflow2.0
-
PyTorch★ 39
miniLLAMA
A compact implementation of the LLaMA and LLaMA 2 architectures for training and inference, written to make the differences from GPT easy to see.
akanyaani/miniLLAMA
-
TensorFlow 2★ 41
ranknet-tensorflow2.0
Learning to rank, from RankNet to LambdaRank, implemented in TensorFlow 2.0.
akanyaani/ranknet-tensorflow2.0
-
TensorFlow★ 7
minGPTF
A TensorFlow re-implementation of Andrej Karpathy's minGPT training code.
akanyaani/minGPTF
Star counts as of October 2026.
Experience
-
May 2026 – Present
FrontiersMind
Co-Founder · Bengaluru
Open foundation models, sovereign and efficient.
-
Sep 2025 – May 2026
Technology Innovation Institute
Senior LLM Research Engineer · Abu Dhabi
Worked on the Falcon LLMs, with a focus on pre-training, model evaluation and post-training refinement.
-
Sep 2024 – Aug 2025
BluOrion
Senior LLM Research Engineer · Dubai
- Designed and led the development of ZClip, an adaptive gradient clipping algorithm that improves stability and eliminates manual batch skipping during LLM training.
- Worked on initialisation and variance control techniques to improve LLM performance.
- Explored parameter-efficient Transformer designs using factorised embeddings and low-rank attention.
- Contributed to distributed pre-training recipe design and built training pipelines on FSDP, PyTorch Lightning and DeepSpeed, used to train 1B to 13B parameter models.
-
Dec 2020 – Sep 2024
yellow.ai
Research Scientist, NLP · Bengaluru
- Co-author of the Komodo LLM.
- Trained and deployed task-specific language models in production.
- Built embedding and generative zero-shot models with distributed pre-training, reducing unidentified utterances by 30%.
- Trained custom GPT and BERT models from scratch on in-house chat data for generation and embeddings.
-
Sep 2017 – Dec 2020
EdGE Networks
Senior Data Scientist, NLP (from Jul 2018) · Bengaluru
Neural ranking and job recommendation models. Implemented a Transformer autoencoder, GPT-2 and a custom autoregressive transformer from scratch in TensorFlow 2.0, with distributed pre-training on a large corpus.
Earlier
- 2016 – 2017Scry Analytics, Data Scientist, Gurgaon. Opinion mining, NER and phrase extraction with LSTMs and CNNs.
- 2015 – 2016Gauge Data Solutions, Data Analyst, Noida. Web crawling and sequence classification for legal documents.
- 2013 – 2014Anyaani Technology, Co-Founder, Noida. Ran Bindaaslo, an online marketplace.
Education
- 2009 – 2014Rajiv Gandhi Prodyogiki Vishwavidyalaya, Bachelor of Technology, Computer Science.
Contact
Working on LLM pre-training, or on serving open models? Write to me at akanyaani@gmail.com.