An open-weight AI model is one whose trained parameters — the billions of numbers that make the model work — can be downloaded and run by anyone, while a regular (closed) model keeps those weights private and only lets you reach it through the company's own app or API.
The practical difference is control: with open weights you can run the model on your own hardware or through a third-party host, inspect it, fine-tune it, and keep using it even if the original maker changes direction. With a closed model, your access lives and dies with the vendor's terms of service.
The mechanism behind this matters more than the label. A model's "weights" are the result of training — a long process where the model reads enormous amounts of text and adjusts its internal numbers to predict what comes next. When a company releases those numbers publicly, it's handing over the finished artifact, not the recipe.
You don't get the training data or the full process, which is why people say "open weight" rather than "open source." Meta's Llama family is the clearest example: according to our AI tool database, Llama is Meta's open-weight mixture-of-experts line, with Llama 4 Maverick at 400B parameters and Scout at 109B, and the weights are a free download.
Mixture-of-experts means the model is split into specialist sub-networks and only routes each request to the relevant ones, which is how a 400B model can still run affordably. But the same database notes Meta shut down the Llama API in July 2026 and has announced a next-generation model called Watermelon — a reminder that "open weights" doesn't mean the project stands still or that the maker's hosted service will stick around.
Here's a concrete worked example. Say you run a small law firm and want a model that drafts summaries of client contracts without sending confidential text to an outside company. With a closed chatbot, every prompt leaves your network.
With an open-weight model like Llama, you download the weights, run them on a machine in your office, and the text never leaves the building. Your costs shift from a monthly subscription to hardware and electricity, and your quality ceiling is whatever the model can do locally. That trade — privacy and control in exchange for setup work and possibly lower raw capability — is the entire reason open weights exist.
Now the limits, because open weights are not automatically better. First, "free download" is not "free to run." A large model needs serious memory and compute; a small team may spend more on infrastructure than a subscription would cost.
Second, hosting varies wildly by provider, so two services running the same Llama model can behave differently in speed and quality. Third, open weights raise governance questions — once released, a model can be fine-tuned for purposes the original creator never intended and cannot recall.
And fourth, open-weight models often lag the newest closed models on hard reasoning tasks, though the gap narrows with each generation. If your work is routine drafting, translation, or classification, an open-weight model is often plenty. If you need the sharpest reasoning on novel problems, a closed frontier model may still win.
Check the vendor's own page for current terms, because licensing and hosting options change frequently.