Text model

LLaMA — open-weight language models

Open LLaMA models in GPTunneL: 1M token context, a free licence and a low price. Pay per token, no subscription and no VPN.

What the LLaMA family is good at

Traits Meta keeps from one generation to the next — you can count on them whichever model you pick.

Open weights

The defining difference: the models are published and can run on your own server. Neither ChatGPT nor Claude offers that.

Low price

Openness creates competition among the providers that run these models — and keeps the price noticeably below closed equivalents.

1M token context

The larger fourth-generation models hold a million tokens — a folder of documents or an entire large project.

The base for fine-tuning

Most open models on the market grew out of LLaMA: it has more ready fine-tunes and tooling around it than anything else.

Which model to pick

The service carries two current fourth-generation models — they differ in size and purpose.

ScoutThe default pickMaverickHard tasks
Best forEveryday writing, correspondence, simple codeHard reasoning, analysis, large documents
Context1M tokens1M tokens
SpeedHighMedium
PriceLowMedium

Figures come from the platform catalogue. Exact per-model prices are on the Pricing

Every LLaMA model

Release dates follow Meta's official announcements.

April 2025

LLaMA 4 ScoutLLaMA 4 Maverick

The first generation with a million tokens of context. Scout is the everyday workhorse, Maverick handles the heavier jobs.

December 2024

LLaMA 3.3 70B

A 70-billion-parameter model that matched the quality of earlier flagships three times its size.

July — September 2024

LLaMA 3.1LLaMA 3.2

Context grew to 128 thousand tokens, very compact versions appeared, and so did the first models that understood images.

April 2024

LLaMA 3

The generation where open models were first seriously compared with closed flagships.

July 2023

LLaMA 2

The first version licensed for commercial use. The whole open-model ecosystem starts here.

Full release history
February 2023

LLaMA 1

The research release that began the history of open large language models.

How billing works

There is no subscription: you top up one balance and spend it on any model on the platform.

Pay per token

You are charged for exactly what the request and the answer used. Not using it costs nothing: no limits, no monthly fee.

One balance for every model

LLaMA, ChatGPT, Claude, image and video generation — all out of the same wallet. No separate subscription per service.

Top up the way you prefer

International cards, Apple Pay, Google Pay or crypto. The minimum top-up is $5.

How to get started with LLaMA

Sign in to GPTunneL
One account for every model on the platform.
Draft an outline for an article about a company moving to remote work
LLaMA 4 Scout
Article outline

Open with the numbers: how many people moved, over what period, and what happened to productivity.

The main section covers the three difficulties they hit and what worked against each one.

Close with a checklist for companies that are only now considering the move.

Sign in whichever way suits you.

The model switches right inside the input — you can change it mid-conversation.

The answer lands in the chat and the context is kept until the conversation ends.

Try LLaMA in GPTunneL

Signing up takes a minute, and $5 on the balance is enough to see whether the model fits your task.

Frequently asked questions

No. Requests go through GPTunneL's infrastructure, so the models open like any ordinary website — no VPN and no proxy.

LLaMA is Meta's family of open neural networks and the foundation of almost the entire open-model ecosystem: most free language models on the market grew out of it in one way or another.

The key difference from ChatGPT and Claude is the published weights. The model can be downloaded and run on your own server, which means nobody can switch it off or change it without warning. Price follows from that openness: anyone can run LLaMA, and the competition keeps the cost well below closed equivalents. The larger fourth-generation models hold a million tokens in context.

LLaMA models are available in GPTunneL without a server of your own — worldwide, without a VPN, with one balance and billing for the tokens you actually spend.