
██ ██ ██ ██ █████ ███ ███
██ ██ ██ ██ ██ ██ ████ ████
███████ ██ ██ ███████ ██ ████ ██
██ ██ ██ ██ ██ ██ ██ ██ ██
██ ██ ██ ▄█ ██ ██ ██ ██ ██
███████ ██ ██ ██████ ███████ ██ ██ █████ ███████ ██ ██ ██ ██ █████ ██████
██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██
███████ ███████ ██████ █████ ████ ███████ ███████ ███████ █████ ███████ ██████
██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██ ██
███████ ██ ██ ██ ██ ███████ ██ ██ ██ ███████ ██ ██ ██ ██ ██ ██ ██ ██● session ready. type for commands.
selected projects
autoexpautoexp
Python
Local-first AI assisted experimentation workspace.
TinyFTTinyFT
Python
TinyFT - Lightweight LLM Fine-Tuning Library
native_sparse_attentionnative_sparse_attention
Python
A pinned project from GitHub.
differential_privacydifferential_privacy
Jupyter Notebook
A pinned project from GitHub.
recent writing
all postsThe only Muon Optimizer guide you need
All neural networks use a form of gradient descent for updating their parameters. The fundamental intuition to all neural net’s parameter optimization seems obvious to us, i.e., to …
A deep-dive into RoPE, and why it matters?
What is Positional Encoding and why it matters? When training any large language model based on Transformers architecture, our input token sequences tend to form a $\text{seq\_len} …
Decoding Karpathy's min-char-rnn (character level Recurrent Neural Network)
Recurrent Neural Networks (RNN) have existed for long at this point, and RNNs without attention mechanism (plain-simple RNN architecture) are no longer the hottest thing either. …
