Meta Pro · $9/month 10 tokens / message Balanced

Llama 3.3 70B · Meta

Meta's Llama 3.3. 70 billion parameters, widely tested in chat, code and reasoning. One of the most solid choices for general use.

Proven and stable Multilingual Good at code

Try it free

Llama 3.3 70B requires the Pro plan

Create a free account to start, or subscribe to a plan for unlimited use of premium models.

No card to start · Cancel anytime

Open in full chat → Compare models side by side, save your sessions and memory

About Llama 3.3 70B

Meta's Llama 3.3. 70 billion parameters, widely tested in chat, code and reasoning. One of the most solid choices for general use.

Where it shines: Proven and stable · Multilingual · Good at code.

How to use Llama 3.3 70B

  1. 1

    Type or upload

    Type what you want in the box above — or upload the file if the tool asks for one.

  2. 2

    Generate

    Click the main button. Wait 2-30 seconds depending on the model and input size.

  3. 3

    Download or share

    Download the result or share the direct link. No watermark, ready to use.

Frequently asked questions

How much does it cost to use Llama 3.3 70B?

Llama 3.3 70B is a Pro model: it costs 10 tokens per use (~$0.05 real cost for us). You need a Pro plan ($9/month → 15,000 tokens) or a one-shot pack. If you already have tokens in the free account, you can also spend them directly.

How many uses of Llama 3.3 70B are included in the Pro plan?

Pro ($9/month) gives you 15,000 recurring tokens. At 10 tokens per use of Llama 3.3 70B, that's ~1,500 full uses per cycle. If you run out, one-shot packs (5,000 / 25,000 / 80,000 tokens) add to the balance without expiring before one year.

What makes Llama 3.3 70B special?

Meta publishes the full weights of the Llama family — well tested in general chat, code and multilingual, with specific strengths in tested and stable, multilingual and good in code.

How fast does Llama 3.3 70B respond?

Llama 3.3 70B has a balanced response time: 5-15 seconds per response — neither the fastest nor the slowest in the catalog. The actual time also depends on the length of the prompt and the load of the datacenter. Models with huge context take longer when you enter very long texts.

How do I use Llama 3.3 70B in ia.gratis?

You can use Call 3.3 70B from /chat/ by selecting Call 3.3 70B in the picker, or via the REST API with `model=call-3` in the body of the POST. Quick summary: versatile and popular. 70B parameters. The internal model identifier is `call-3` — useful when integrating via API.