Abhay Kumar

LLM pre-training · Bengaluru, India

Abhay Kumar

I work on how large language models are trained: stability, optimisation and scale.

Co-founder, FrontiersMind, building open foundation models that are sovereign and efficient. Author of ZClip.

Schematic of a gradient-norm trace with spikes clipped to an adaptive threshold
raw gradient norm adaptive threshold clipped Schematic, not experimental data: the idea behind ZClip.
01

About

I am an LLM researcher and the co-founder of FrontiersMind. My work is on pre-training: what makes large runs diverge, how to initialise and control variance so they do not, and how to build the distributed pipelines that train them.

Before FrontiersMind I was a Senior LLM Research Engineer at the Technology Innovation Institute in Abu Dhabi, working on the Falcon LLMs, and at BluOrion in Dubai, where I designed ZClip, an adaptive gradient clipping algorithm that removes the need to skip batches by hand, and worked on initialisation and variance control. I have co-led the pre-training of 7B and 13B parameter models on 15 trillion tokens. Earlier, at yellow.ai, I co-authored the Komodo LLM for Indonesia's regional languages.

I have worked in machine learning and NLP for more than ten years. Several from-scratch implementations I wrote along the way, of GPT-2 in TensorFlow and of LLaMA in PyTorch, are open source.

7B · 13B
parameter models whose pre-training I co-led
15T
tokens in that pre-training run
FSDP · DeepSpeed · Lightning
distributed training stacks I build pipelines on
02

Research

Google Scholar

  1. 2025arXiv preprintFirst author

    ZClip: Adaptive Spike Mitigation for LLM Pre-Training

    Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury, Fabian Güra

    An adaptive gradient clipping algorithm that sets its threshold from running statistics of the gradient norm, using z-score anomaly detection to catch the spikes that cause loss blow-ups without interfering with convergence otherwise.

  2. 2025arXiv preprint

    A Refined Analysis of Massive Activations in LLMs

    Louis Owen, Nilabhra Roy Chowdhury, Abhay Kumar, Fabian Güra

    An analysis of massive activations across a broad range of LLM architectures. Not all of them are detrimental, and known mitigations are model-specific; pairing Target Variance Rescaling with Attention KV bias or Dynamic Tanh balances mitigation with downstream performance.

  3. 2025arXiv preprint

    Variance Control via Weight Rescaling in LLM Pre-training

    Louis Owen, Abhay Kumar, Nilabhra Roy Chowdhury, Fabian Güra

    Introduces Layer Index Rescaling for weight initialisation and Target Variance Rescaling for variance control. On a 1B parameter LLaMA model, better variance management improves downstream performance by up to 4.6% and reduces extreme activation values.

  4. 2024arXiv preprint

    Komodo: A Linguistic Expedition into Indonesia's Regional Languages

    Louis Owen, Vishesh Tripathi, Abhay Kumar, Biddwan Ahmed

    Komodo-7B, a family of 7-billion-parameter LLMs that work across Indonesian, English and 11 regional languages of Indonesia.

  5. 2023ICAICTA 2023 (IEEE)

    BED: Bi-Encoder-Based Detectors for Out-of-Distribution Detection

    Louis Owen, Biddwan Ahmed, Abhay Kumar

    Bi-encoder-based detectors for out-of-distribution detection in NLP, with a comparison of OOD methods and feature extractors on CLINC150, ROSTD-Coarse, SNIPS and YELLOW.

03

Open source

All repositories

  • PythonNew, 2026

    faithserve

    Is this endpoint serving this open-weight model faithfully? A CLI that runs black-box checks against an OpenAI-compatible server: chat template parity, sampling parameters, streaming and tool calls.

    FrontiersMindAI/faithserve

  • PyTorch★ 153

    ZClip

    Official implementation of the ZClip paper: adaptive gradient clipping from EMA statistics of the gradient norm, with optional PyTorch Lightning support.

    bluorion-com/ZClip

  • TensorFlow 2★ 267

    gpt-2-tensorflow2.0

    OpenAI GPT-2 pre-training and sequence prediction, implemented in TensorFlow 2.0.

    akanyaani/gpt-2-tensorflow2.0

  • PyTorch★ 39

    miniLLAMA

    A compact implementation of the LLaMA and LLaMA 2 architectures for training and inference, written to make the differences from GPT easy to see.

    akanyaani/miniLLAMA

  • TensorFlow 2★ 41

    ranknet-tensorflow2.0

    Learning to rank, from RankNet to LambdaRank, implemented in TensorFlow 2.0.

    akanyaani/ranknet-tensorflow2.0

  • TensorFlow★ 7

    minGPTF

    A TensorFlow re-implementation of Andrej Karpathy's minGPT training code.

    akanyaani/minGPTF

Star counts as of October 2026.

04

Experience

  1. May 2026 – Present

    FrontiersMind

    Co-Founder · Bengaluru

    Open foundation models, sovereign and efficient.

  2. Sep 2025 – May 2026

    Technology Innovation Institute

    Senior LLM Research Engineer · Abu Dhabi

    Worked on the Falcon LLMs, with a focus on pre-training, model evaluation and post-training refinement.

  3. Sep 2024 – Aug 2025

    BluOrion

    Senior LLM Research Engineer · Dubai

    • Designed and led the development of ZClip, an adaptive gradient clipping algorithm that improves stability and eliminates manual batch skipping during LLM training.
    • Worked on initialisation and variance control techniques to improve LLM performance.
    • Explored parameter-efficient Transformer designs using factorised embeddings and low-rank attention.
    • Contributed to distributed pre-training recipe design and built training pipelines on FSDP, PyTorch Lightning and DeepSpeed, used to train 1B to 13B parameter models.
  4. Dec 2020 – Sep 2024

    yellow.ai

    Research Scientist, NLP · Bengaluru

    • Co-author of the Komodo LLM.
    • Trained and deployed task-specific language models in production.
    • Built embedding and generative zero-shot models with distributed pre-training, reducing unidentified utterances by 30%.
    • Trained custom GPT and BERT models from scratch on in-house chat data for generation and embeddings.
  5. Sep 2017 – Dec 2020

    EdGE Networks

    Senior Data Scientist, NLP (from Jul 2018) · Bengaluru

    Neural ranking and job recommendation models. Implemented a Transformer autoencoder, GPT-2 and a custom autoregressive transformer from scratch in TensorFlow 2.0, with distributed pre-training on a large corpus.

Earlier

  • 2016 – 2017Scry Analytics, Data Scientist, Gurgaon. Opinion mining, NER and phrase extraction with LSTMs and CNNs.
  • 2015 – 2016Gauge Data Solutions, Data Analyst, Noida. Web crawling and sequence classification for legal documents.
  • 2013 – 2014Anyaani Technology, Co-Founder, Noida. Ran Bindaaslo, an online marketplace.
05

Education

  • 2009 – 2014Rajiv Gandhi Prodyogiki Vishwavidyalaya, Bachelor of Technology, Computer Science.
06

Contact

Working on LLM pre-training, or on serving open models? Write to me at akanyaani@gmail.com.