Skip to main content
Built for real engineering decisions

AI Model & API Guides

Practical tutorials · Cost boundaries · Troubleshooting

Source-linked guidance for AI models, API integration, and developer workflows—turning fast-changing prices, quotas, capabilities, and limits into decisions you can verify.

  • Primary sources first
  • Reproducible steps and failure boundaries
  • Volatile facts carry review dates
440
Localized guides
6
Language markets
Sep 6, 2026
Latest update

440 articles

Claude Sonnet 5 pricing after the canceled September increaseClaude

Claude Sonnet 5 Price Increase Canceled: Current Rates and Billing Math

Anthropic canceled the September 1 Sonnet 5 API price increase. See the current USD rates, what changes in a September budget, and how to reconcile uncached input, cache writes, cache reads, and output.

FLUX.2 Pro and Flex illustrated as production and adjustable-detail workflowsAI Image Generation

FLUX.2 Flex vs Pro: Which Model Is Worth Using?

Start with FLUX.2 Pro for lower generation costs. Test Flex when text or fine details need more control. Compare current BFL pricing, editing costs, and the acceptance rate that would justify paying more.

Choosing a single RTX 5090 or dual GPUs for Qwen3.8-27B based on precision, context, and concurrency.AI Models

Qwen3.8-27B VRAM Guide: RTX 5090, Dual GPUs, and 262K Context

A 32GB RTX 5090 can run quantized Qwen3.8-27B, but fitting 262,144 tokens depends on the weights, KV cache, runtime, and workload. Use these memory budgets and configuration checks to choose a practical single- or dual-GPU setup.

A request passing through an optional API Management policy and Azure OpenAI deployment checks that can return 429.API

Azure OpenAI TPM Rate Limits: Diagnose 429s Below Quota

Azure OpenAI can return 429 while billed token usage looks low. Compare the receiving deployment’s allocation with response headers, account for estimated tokens and request bursts, and verify recovery without hiding failures behind retries.

GPT-6 Astra request workflow linking the service, project key, and model to response verificationAI API

GPT-6 Astra API Access: Check Your Project and Make a First Request

GPT-6 Astra API access depends on the service, organization, and project behind your key. Use a complete first-request example to check access, read the response correctly, and identify the next step when a call fails.

Astra limits separated into Chat message caps, the shared Work and Codex allowance, and API billing and throughputAI Development Tools

Codex Usage Limits: Check Your Astra Allowance and Get Back to Work

Check your Codex allowance and reset time, then decide whether to wait, use a banked reset, or buy more usage. GPT-6 Astra has different limits in Chat, Work, and the API.

Sora 2 API removal scheduled for September 24, 2026, separate from the April 26 Sora app discontinuationAI Video Generation

Sora 2 API shutdown: the September 24 deadline and what to do now

OpenAI has scheduled the Sora 2 models and Videos API for removal on September 24, 2026. Here is how to separate the app, direct API, and Microsoft Foundry deadlines—and preserve your work before switching providers.

Sora 2 API migration to Veo 3.1 with the scheduled September 24, 2026 removal date.AI Video Generation

Sora 2 vs. Veo 3.1 API: Costs, Compatibility, and Migration

Veo 3.1 can replace short-video generation in a Sora integration, but requests, job tracking, assets, and output constraints change. Compare current API costs and follow the migration through to a saved video.

Illustration for creating and checking transparent PNG assetsAI Image Editing

Make a Transparent PNG with GPT Image 2—or Remove an Existing Background

A reusable cutout needs more than a .png extension. Start with the right creation method, then check the downloaded file and the background where it will actually appear.

Gemini 3.8 Flash and Claude Fable 5.1 API decision visual for developersAI Model Comparison

Gemini 3.8 Flash vs Claude Fable 5.1: Which API Should You Test First?

Gemini 3.8 Flash is the lower-cost first test for most high-volume workloads. Claude Fable 5.1 is a targeted challenger for demanding long-running work—but only your own acceptance, latency, and repair data can justify the premium.

Workload decision board comparing GPT-6 Astra and Claude Fable 5.1 across access, API cost, data rules, and acceptance testingAI Model Comparison

GPT-6 Astra vs Claude Fable 5.1: Choose by Workload

Astra and Fable share the same headline input and output rates, but cache pricing, long-context billing, access, and safeguards can reverse the cheaper or safer choice.

Dashboard for estimating an Astra API pilot from token rates, the 272K price band, add-ons, and accepted-task costAI Models

GPT-6 Astra API Pricing: Calculate Your Real Pilot Cost

The $10/$50 headline is only the short-context Standard rate. Price cache writes, full-request long-context repricing, service tiers, tools, and failed attempts before a pilot.

GPT-6 Astra computer-use overview showing integration choices, the Responses API loop, permission checks, state tracking, and outcome verificationAI Development Tools

GPT-6 Astra Computer Use: Build the Right Responses API Loop

GPT-6 Astra can plan multi-step work across browser and desktop interfaces, but your application still owns execution, permission checks, persistent state, and proof that the external task actually succeeded.

Technical illustration for GPT-6 Astra API rate limits, rollout access, and 429/503 recovery.API Guides

OpenAI API Rate Limits: GPT-6 Astra Tiers and 429/503 Recovery

GPT-6 Astra has a public tier table, but your usable capacity still depends on rollout access, organization and project limits, and the response that actually failed.

Claude Free and Pro usage decision based on five-hour and weekly countersClaude

Claude Free vs Pro Usage Limits: Decide With Work, Not Message Counts

Pro gives at least five times Free usage per session during peak periods, but neither plan promises a fixed message count. Use your own interruptions to price the upgrade.

Claude Fable 5.1 and Opus 5 cost decision separating uncached input, cache writes, cache reads, output, and accepted tasksAPI Guides

Claude Fable 5.1 vs Opus 5: Pricing and Cost per Task

Opus 5 costs half as much for uncached input and output. Fable 5.1 costs half as much for cache reads. Your workload and acceptance rate decide which total is lower.

Claude Fable 5.1 plan limits for Pro and Max usersClaude

Claude Fable 5.1 Limits: Pro, Max, and the 50% Cap

Max includes Fable 5.1 within a shared weekly pool, capped at 50% for Fable models. Pro uses paid usage credits from the first Fable request.

Document cards moving through separate Claude chat, Project, context, and usage limit gatesClaude

Does Claude Have Unlimited File Uploads? The Limits That Actually Apply

Claude Projects can hold an unlimited number of files, but that narrow statement comes with per-file and context conditions. Chats, PDFs, images, and account usage each have their own ceilings.

Claude Free and Pro decision with a monthly cost and interruption logClaude

Is Claude Pro Worth It in 2026? A Practical $20 Decision

Claude Pro is worth it when Free repeatedly breaks a task you value or when a paid-only capability becomes part of your weekly work. If neither is true, Free is the better plan.

DeepSeek V4 peak and off-peak schedule, token rates, cost math, and workload decision overviewAI API Guides

DeepSeek V4 Peak Pricing: Calculate the Real Cost Before You Reschedule

DeepSeek charges peak rates during two weekday UTC windows and half-price rates at all other times. The discount matters only when a workload can move without missing its deadline or adding failures.

Gemini 3.8 Flash Cyber and Fairwind controlled-access conceptAI Security

Gemini 3.8 Flash Cyber Access: A Fairwind Eligibility Guide

Fairwind is an organization-vetted cyber-defense program, not a public API tier. Use this guide to decide whether applying is realistic and what must be ready first.

Decision path from Gemini 3.8 Flash pricing and a 3.7 canary to the restricted Flash Cyber programAI Models

Gemini 3.8 Flash: Pricing, Model ID, and a Safe 3.7 Migration

Gemini 3.8 Flash is ready to evaluate, but its 3.7-matching token rates do not guarantee the same bill. Measure accepted work, token use, latency, and tools—and treat Cyber as a separate Fairwind application.

GLM-5.3-Flash selection guide covering Ox Alpha identity, API pricing, a hosted call, and local memory capacityAI Model Guides

GLM-5.3-Flash Guide: Ox Alpha, API Pricing, and Local Deployment

Start with the hosted API unless your GPUs can hold roughly 305.8 GiB of FP8 weights plus runtime and cache headroom. The 18B active-parameter figure is not a memory requirement.

Overview of Grok Bot's persistent cloud computer, shared access, approval points, and a low-risk first taskAI Tools

Grok Bot's Cloud Computer: What Persists, What Is Shared, and What to Approve

The feature that keeps Grok Bot working after you close your laptop also keeps sessions and files available. Learn the account-wide boundary before handing it email, admin access, or production work.

Showing 24 / 440