NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Instructions to use TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoTokenizer, AutoModelForMultimodalLM

tokenizer = AutoTokenizer.from_pretrained("TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True)
model = AutoModelForMultimodalLM.from_pretrained("TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Notebooks
Google Colab
Kaggle
Local Apps Settings

vLLM

How to use TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

SGLang

How to use TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Docker Model Runner:
```
docker model run hf.co/TitanML/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
```

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 / bias.md

jamesdborin

Duplicate from nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

6c06e7e 2 months ago

preview code

raw

history blame contribute delete

2.63 kB

Field	Response
Participation considerations from adversely impacted groups protected classes in model design and testing:	None
Bias Metric (If Measured):	BBQ Accuracy Scores in Ambiguous Contexts
Which characteristic (feature) show(s) the greatest difference in performance?:	The model shows high variance in the characteristics when it is used with a high temperature.
Which feature(s) have the worst performance overall?	Physical Appearance
Measures taken to mitigate against unwanted bias:	Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) employed to calibrate the model’s reasoning capabilities to maintain logical consistency and appropriate complexity when interacting with or interpreting data from diverse age demographics.
If using internal data, description of methods implemented in data acquisition or processing, if any, to address the prevalence of identifiable biases in the training, testing, and validation data:	The training datasets contain a large amount of synthetic data generated by LLMs. We manually curated prompts.
Tools used to assess statistical imbalances and highlight patterns that may introduce bias into AI models:	BBQ
Tools used to assess statistical imbalances and highlight patterns that may introduce bias into AI models:	These datasets, such as web-scraped finance reasoning data, do not collectively or exhaustively represent all demographic groups (and proportionally therein). For instance, these datasets do not contain explicit mentions of the following classes: age, gender, or ethnicity in approximately 97% to 99% of samples. Finance reasoning data scraped from SEC EDGAR contained a notable representational skew where ethnicity mentions are dominated by Middle Eastern contexts (found in finance documents), while gender is explicitly mentioned in only 0.9% of samples (including Male-only, Female-only, and Both). To mitigate these imbalances, we recommend considering these evaluation techniques such as bias audits, fine-tuning with demographically balanced datasets, and mitigation strategies such as counterfactual data augmentation to align with the desired model behavior. This evaluation used a 3,000-sample subset per dataset, identified as the optimal threshold for maximizing embedder accuracy.
Unwanted Bias Testing:	Constrained to English-language inputs. Multi-lingual parity is not currently claimed or guaranteed.