NVIDIA Pro · $9/month 10 tokens / message Balanced

NVIDIA: Nemotron 3.5 Lightning · NVIDIA

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 3B active parameters out of 30B total. It's optimized for high-throughput agentic workloads and specialized tasks requiring fast and accurate responses.

High throughput Expert architecture Fast responses

Try it free

NVIDIA: Nemotron 3.5 Lightning requires the Pro plan

Create a free account to start, or subscribe to a plan for unlimited use of premium models.

No card to start · Cancel anytime

Open in full chat → Compare models side by side, save your sessions and memory

About NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 3B active parameters out of 30B total. It's optimized for high-throughput agentic workloads and specialized tasks requiring fast and accurate responses.

Where it shines: High throughput · Expert architecture · Fast responses.

How to use NVIDIA: Nemotron 3.5 Lightning

  1. 1

    Type or upload

    Type what you want in the box above — or upload the file if the tool asks for one.

  2. 2

    Generate

    Click the main button. Wait 2-30 seconds depending on the model and input size.

  3. 3

    Download or share

    Download the result or share the direct link. No watermark, ready to use.

Frequently asked questions

How much does it cost to use NVIDIA: Nemotron 3.5 Lightning?

NVIDIA: Nemotron 3.5 Lightning is a Pro model: it costs 10 tokens per use (~$0.05 real cost for us). You need a Pro plan ($9/month → 15,000 tokens) or a one-shot pack. If you already have tokens in the free account, you can also spend them directly.

How many uses of NVIDIA: Nemotron 3.5 Lightning are included in the Pro plan?

Pro ($9/month) gives you 15,000 recurring tokens. At 10 tokens per use of NVIDIA: Nemotron 3.5 Lightning, that's ~1,500 full uses per cycle. If you run out, one-shot packs (5,000 / 25,000 / 80,000 tokens) add to the balance without expiring before one year.

What makes NVIDIA: Nemotron 3.5 Lightning special?

NVIDIA fine-tune its models for fast inference on its own optimized hardware — good at technical questions and reasoning, with specific strengths in high performance, expert architecture, and quick response.

How fast does NVIDIA: Nemotron 3.5 Lightning respond?

NVIDIA: Nemotron 3.5 Lightning has a balanced speed: 5-15 seconds per response — neither the fastest nor the slowest in the catalog. The actual time also depends on the length of the prompt and the load of the datacenter — models with huge context take longer when you enter very long texts.

How do I use NVIDIA: Nemotron 3.5 Lightning in ia.gratis?

You can use NVIDIA: Nemotron 3.5 Lightning from /chat/ by selecting NVIDIA: Nemotron 3.5 Lightning in the picker, or via the REST API with `model=nemotron-3-5-lightning-2` in the POST body. Quick summary: nVIDIA Nemotron 3.5 Lightning: Expert model of 3B active parameters for agile tasks. The internal model identifier is `nemotron-3-5-lightning-2` — useful when integrating by API.