Page cover
For the complete documentation index, see llms.txt. This page is also available as Markdown.

Arcee Model Overview

Trinity-Large-Thinking is the only Arcee model available via api. See full list of available open models on pricing.

Model
Trinity-Nano (6B)
Trinity-Mini (26B)
Trinity-Large-Thinking (400B)

Strength

Lightweight, ultra-low latency model.

Fast and cost-efficient model for well-defined tasks.

Robust generalist model with strong performance across reasoning, coding, math, and complex task decomposition.

Ideal Deployment

Fully local on consumer GPUs, edge servers, and mobile devices. Tuned for offline operation.

Serve customer-facing apps, agent backends, and high-throughput services in cloud or VPC.

Advanced agents, reasoning systems, and developer tools. Deployed via hosted cloud endpoints or self-hosted in multi-GPU configurations.

Active Parameters

1B per token

3B per token

13B per token

Context Window

128k tokens

128k tokens

512k tokens (hosted at 128k)

Knowledge Cutoff

2024

2024

2024

Speed

Instant

Very Fast

Very Fast

API Model Name

Not Hosted

Not Hosted

trinity-large-thinking

Last updated