# CARouter > CARouter is an OpenAI-compatible API gateway in front of AI providers that run their inference on Canadian soil. One base URL, one key, and the choice of which companies - in which province, under which ownership - may see your prompts. Billed in CAD. ## How to call it ```python from openai import OpenAI client = OpenAI( base_url="https://carouter.ai/v1", api_key="car_...", # from https://carouter.ai/dashboard/keys ) response = client.chat.completions.create( model="qwen3.6-35b-a3b", messages=[{"role": "user", "content": "Bonjour"}], ) ``` Endpoints: `POST /v1/chat/completions`, `POST /v1/completions`, `POST /v1/responses`, `POST /v1/embeddings`, `POST /v1/rerank`, `POST /v1/ner`, `POST /v1/moderations`, `POST /v1/decisions`, `GET /v1/models`. Streaming with `"stream": true` on `/v1/chat/completions`, `/v1/completions`, `/v1/responses`; every other endpoint has no streaming form and ignores the flag. Errors use OpenAI's shape: a top-level `{"error": {...}}`. ## Documentation - [CARouter for language models](https://carouter.ai/llms-full.txt): every endpoint, every model, every price and every error code, in one plain-text fetch. Read this one. - [Documentation](https://carouter.ai/docs): the same reference as a web page, with live prices per provider. - [Model catalog JSON](https://carouter.ai/api/public/models): the catalog this file is generated from. Takes `?plan=free|pro|scale`. - [Tools](https://carouter.ai/docs/tools.md): the middlewares an API key can enable, one line each, with a link to each tool's own page. - [Provider network](https://carouter.ai/#providers): who serves traffic, where, and who owns them. ## Models (25 routable) Prices are list prices on the Free plan (15% over the provider's base rate), in CAD per million tokens. - `tofino-3`: 256K context; Proprietary (Augure); vision; tools; json schema; 0.69/2.30 CAD per M tokens in/out; routable now - `deepseek-v4-flash-0731`: 256K context; MIT; reasoning; tools; json schema; 0.7401/1.4802 CAD per M tokens in/out; routable now - `rosedale-1`: 1M context; Proprietary (Augure); reasoning; vision; tools; json schema; 2.875/6.90 CAD per M tokens in/out; routable now - `glm-5.2`: 256K context; MIT; reasoning; tools; json schema; 0.4776/1.5122 CAD per M tokens in/out; routable now - `mistral-medium-3.5-128b`: 128B; 176K context; Modified MIT; reasoning; vision; tools; json schema; 2.7754/13.8768 CAD per M tokens in/out; routable now - `qwen3.6-27b`: 27B; 256K context; Apache 2.0; reasoning; vision; tools; json schema; 0.7401/4.9956 CAD per M tokens in/out; routable now - `qwen3.6-35b-a3b`: 35B; 256K context; Apache 2.0; reasoning; vision; tools; json schema; 0.23/0.92 CAD per M tokens in/out; routable now - `gemma-4-26b-a4b-it`: 26B; 256K context; Apache 2.0; reasoning; vision; tools; json schema; 0.4625/0.9252 CAD per M tokens in/out; routable now - `qwen3.5-9b`: 9.7B; 256K context; Apache 2.0; reasoning; vision; tools; json schema; 0.185/0.2775 CAD per M tokens in/out; routable now - `qwen3.5-397b-a17b`: 397B; 244K context; Apache 2.0; reasoning; vision; tools; json schema; 1.1101/6.6608 CAD per M tokens in/out; routable now - `gpt-oss-20b`: 21B; 128K context; Apache 2.0; reasoning; tools; json schema; 0.0741/0.2775 CAD per M tokens in/out; routable now - `qwen3-coder-30b-a3b-instruct`: 30B; 128K context; Apache 2.0; tools; json schema; 0.3701/1.4802 CAD per M tokens in/out; routable now - `qwen3-235b-a22b-instruct-2507`: 235B; 244K context; Apache 2.0; tools; json schema; 0.1433/0.8755 CAD per M tokens in/out; routable now - `mistral-small-3.2-24b-instruct-2506`: 24B; 128K context; Apache 2.0; vision; tools; json schema; 0.2775/0.6476 CAD per M tokens in/out; routable now - `qwen3-0.6b`: 0.6B; 40K context; Apache 2.0; reasoning; 0.115/0.46 CAD per M tokens in/out; routable now - `qwen2.5-vl-72b-instruct`: 72B; 32K context; Qwen; vision; json schema; 1.6837/1.6837 CAD per M tokens in/out; routable now - `pixtral-12b-2409`: 12B; 128K context; Apache 2.0; vision; tools; json schema; 0.3701/0.3701 CAD per M tokens in/out; routable now - `kev-0.8b`: 0.8B; 8K context; Apache-2.0; 0.0306/0.00 CAD per M tokens in/out; routable now - `kev-4b`: 4B; 8K context; Apache-2.0; 0.0612/0.00 CAD per M tokens in/out; routable now - `qwen3-embedding-8b`: 7.6B; 32K context; Apache 2.0; 0.115/0.00 CAD per M tokens in/out; routable now - `bge-multilingual-gemma2`: 8K context; Gemma; 0.0185/0.00 CAD per M tokens in/out; routable now - `bge-m3`: 0.567B; 8K context; MIT; 0.0185/0.00 CAD per M tokens in/out; routable now - `qwen3guard-gen-0.6b`: 0.6B; 32K context; Apache 2.0; 0.115/0.00 CAD per M tokens in/out; routable now - `carouter-ner-multilingual`: 195K context; MIT; 0.115/0.00 CAD per M tokens in/out; routable now - `bge-reranker-v2-m3`: 0.568B; 8K context; Apache-2.0; 0.0172/0.00 CAD per M tokens in/out; routable now ## Providers (4 live, 5 announced) A provider that is not both Canadian-owned and Canadian-hosted is switched OFF for new accounts and stays off until the account holder turns it on. - Augure (Canada; Canadian-owned; zero data retention; live; off by default) - CARouter Cloud (Ottawa, ON; Canadian-owned; zero data retention; live; on by default) - SPUR Compute (Waterloo, ON; Canadian-owned; limited retention; announced, not routable yet; on by default) - PolarGrid (Canada; Canadian-owned; retention not verified; announced, not routable yet; off by default) - TELUS Sovereign AI (Rimouski, QC; Canadian-owned; retention not verified; contacted, not routable yet; on by default) - Bell AI Fabric (Kamloops, BC; Canadian-owned; retention not verified; contacted, not routable yet; on by default) - Cohere (Iowa (GCP us-central1), United States; Canadian-owned; retention not verified; contacted, not routable yet; off by default) - Scaleway (Paris, France; foreign-owned; zero data retention; live; off by default) - OVHcloud AI Endpoints (Gravelines, France; foreign-owned; zero data retention; live; off by default) ## Notes - CARouter is in alpha. The gateway, metering and billing are real; the provider network is still being signed up, which is why providers are marked live or announced above. Every model listed is routable today: a model nobody serves yet is not in this file. - Prompts and completions are not stored. Request metadata (model, provider, token counts, latency, cost) is, because billing and usage history are made of it. - Generated 2026-09-30 from the live catalog.