AI infrastructure startup SubQ, co-founded by Alexander Whedon (who also serves as CTO), today announced early access to its SubQ model — a large language model built on a fully sub-quadratic sparse attention architecture (SSA). The company also launched SubQ Code, a coding agent. SubQ claims this is the first frontier-level model that natively supports a 12 million token context window.
Traditional Transformer models compute attention between every pair of tokens, causing compute to grow quadratically with input length. SubQ's architecture extracts only a few critical connections, dramatically cutting the workload. At a 12 million token scale, the company says compute drops by nearly 1,000x; at 1 million tokens, inference speed is 52 times faster than FlashAttention. The API is compatible with OpenAI's format and costs just $0.08 per million tokens.
SubQ's benchmark claims are ambitious. On the MRCR V2 test, designed to evaluate across 1 million token contexts, SubQ scored 83% accuracy, beating Claude Opus 4.6's 76% and GPT-5.4's roughly 36%. On SWE Bench Verified, a practical coding benchmark, SubQ claims 82.1%, again ahead of the same rivals. And in the "single needle in a haystack" test at 12 million tokens, the model maintained 92% accuracy.