- Falcon-H1
- Worked on TII's hybrid Transformer and Mamba LLM family and the Falcon coder models.
- ZClip
- First author. Adaptive gradient clipping that stops loss spikes in LLM pre-training.
- 24B+ models
- Hands-on LLM pre-training experience, up to models of 24B+ parameters.
- GPT-2 in 2019
- One of the first open-source GPT-2 implementations in TensorFlow 2.0. 267 GitHub stars, 78 forks.
-
TensorFlow 2★ 267
gpt-2-tensorflow2.0
One of the first GPT-2 implementations in TensorFlow 2.0, open-sourced in August 2019, before TensorFlow 2.0's stable release. Pre-training and text generation from scratch, with distributed training.
akanyaani/gpt-2-tensorflow2.0
-
PyTorch★ 153
ZClip
Official implementation of the ZClip paper: adaptive gradient clipping from EMA statistics of the gradient norm, with optional PyTorch Lightning support.
bluorion-com/ZClip
-
TensorFlow 2★ 41
ranknet-tensorflow2.0
Learning to rank, from RankNet to LambdaRank, implemented in TensorFlow 2.0.
akanyaani/ranknet-tensorflow2.0
-
PyTorch★ 39
miniLLAMA
A compact implementation of the LLaMA and LLaMA 2 architectures for training and inference, written to make the differences from GPT easy to see.
akanyaani/miniLLAMA
Star counts as of October 2026.
Experience
-
Sep 2025 – May 2026
Technology Innovation Institute
Senior LLM Research Engineer · Abu Dhabi
Worked on the Falcon-H1 models, TII's hybrid Transformer and Mamba LLM family, and on the Falcon coder models, across pre-training, evaluation and post-training refinement.
Evaluation at scale. Designed and built a scalable evaluation pipeline for the coder models from scratch, on Kubernetes and Docker.
-
Sep 2024 – Aug 2025
BluOrion
Senior LLM Research Engineer · Dubai
Designed and led ZClip. Worked on initialisation and variance control, and built distributed pre-training pipelines on FSDP, PyTorch Lightning and DeepSpeed, used to train 1B to 13B parameter models.
-
Dec 2020 – Sep 2024
yellow.ai
Research Scientist, NLP · Bengaluru
Co-authored the Komodo LLM. Trained and deployed task-specific language models in production, and built zero-shot embedding and generative models that reduced unidentified utterances by 30%.
-
Sep 2017 – Dec 2020
EdGE Networks
Senior Data Scientist, NLP · Bengaluru
Early GPT pre-training. In 2019 I implemented GPT-2, a Transformer autoencoder and a custom autoregressive transformer from scratch in TensorFlow 2.0, and pre-trained them with distributed training on a large corpus.
Open-sourced GPT-2 in 2019. I released it as gpt-2-tensorflow2.0, one of the first GPT-2 implementations in TensorFlow 2.0, published before TensorFlow 2.0's stable release. It now has 267 GitHub stars and 78 forks.
Also built neural ranking and job recommendation models.
-
May 2016 – Sep 2017
Scry Analytics
Data Scientist · Gurgaon
Built opinion mining, named-entity recognition and phrase extraction models with LSTMs and CNNs, and a PySpark pipeline to run them on large document sets.
-
May 2015 – May 2016
Gauge Data Solutions
Data Analyst · Noida
Built a multithreaded web crawler for legal documents and a sequence classification model to clean the crawled data.
Education
- 2009 – 2014Rajiv Gandhi Prodyogiki Vishwavidyalaya, Bachelor of Technology, Computer Science.
Contact
Working on LLM pre-training? Write to me at akanyaani@gmail.com.