New chat

Trained from scratch.
Ask it anything.

A 1.32B-parameter Mixture-of-Experts model pretrained on 18B tokens, then tuned with SFT and DPO. Only 0.28B parameters run per token. It writes and rewrites well; for anything factual, upload a document and it will answer from that.

Trained on 18B tokens — about 500x less than comparable 1B models. It writes fluently but does not reliably know facts.