In partnership with

Intelligence Brief

A Chinese AI lab just handed away the biggest open model on record. Moonshot AI released Kimi K3 in full, model weights, technical report, and infrastructure stack together, ahead of their own July 27 deadline. At 1.4 terabytes in size, it's the largest open weight model anyone has shipped. Anyone with the right hardware can download and run it without paying for an API.

We covered Kimi K3 last week too, back when Microsoft was testing it inside Copilot and Elon Musk praised it in an interview. This release makes it real for developers everywhere, not just inside a lab.

The integrated coworker for AI native teams

Empower your team to do their best work with Adapt, the integrated coworker that works alongside your team in Slack and deeply understands your business.

Here’s how Adapt is different

  • Set up takes minutes: connect your tools, add to Slack, and it’s right there for anyone to tag @Adapt for help

  • Does real, high-ROI work: automates work on a schedule; builds internal tools with live data; and does complex, multi-tool tasks on demand

  • Learns your business as you work, becomes your company brain

  • Uses the best AI model for the task, not tied to a single provider

  • SOC2 Type II, RBAC, and support for personal and company-wide integrations

Story Breakdown

The numbers are big. 2.8 trillion total parameters, but only 104 billion active per token, spread across 896 experts with 16 firing at once. That mixture of experts design is what makes a model this size usable at all.

Context window runs to 1 million tokens. The model handles images and video natively too. Two new architecture pieces, Kimi Delta Attention and Attention Residuals, push intelligence per unit of compute about 2.5 times higher than the earlier Kimi K2.

Pricing works out to $0.30 per million tokens for cache hit input, $3 for cache miss input, and $15 per million output tokens.

Moonshot backed the release with case studies that are hard to ignore. Given 48 hours, the model designed, optimized, and verified its own 4 square millimeter chip, one capable of processing over 8,700 tokens per second in simulation. It also built a GPU compiler called MiniTriton from scratch, and in some benchmarks it beat the industry standard Triton compiler outright.

One more example stands out. An astrophysics research task that would normally take an experienced researcher one to two weeks, the model finished in two hours, reviewing more than 20 papers and checking over 300 equations along the way.

Own Search With Podcasts

Search engines and AI platforms reward brands that are mentioned, cited, and trusted across the web. PodPitch books your experts on relevant shows, creating branded mentions, backlinks, transcripts, citations, and reusable content. Only 20 demo spots are available this month. Once claimed, the offer disappears.

Strategic Perspective

Moonshot didn't hide the rough edges either. Their own technical report admits K3 still trails Claude Fable 5 and GPT 5.6 Sol on user experience, even while calling it highly competitive overall. They also flagged a real risk, the model can get overly proactive on complex tasks and start making decisions the user never asked for.

Running this thing is its own barrier. A 2.8 trillion parameter model needs a supernode setup of 64 or more accelerators at minimum. That's not something most companies, in Bangladesh or anywhere else, can just spin up on their own.

So is this the democratization Moonshot wants us to see, or does it mostly hand an advantage to whoever already owns the compute? The model is free. The hardware to run it well is not. That gap is where the real story sits.

Your Time Tool Doesn't Talk To Your Billing

So someone reconciles it by hand, every month. Timeglass sends approved time straight into QuickBooks, Gusto, ADP, Excel, or Sheets. No copy-paste. See the integrations.

Keep Reading