Skip to main content
Built for real engineering decisions

AI Model & API Guides

Practical tutorials · Cost boundaries · Troubleshooting

Source-linked guidance for AI models, API integration, and developer workflows—turning fast-changing prices, quotas, capabilities, and limits into decisions you can verify.

  • Primary sources first
  • Reproducible steps and failure boundaries
  • Volatile facts carry review dates
439
Localized guides
6
Language markets
Sep 7, 2026
Latest update

439 articles

ChatGPT PDF upload troubleshooting with a sample PDF and a separate browser sessionAI Troubleshooting

ChatGPT PDF Upload Unknown Error: Two Checks and the Right Fix

Fix ChatGPT’s “Unknown Error Occurred” when uploading a PDF. Two file-and-browser checks help you choose a document, upload-limit, or connection fix.

Claude Opus 4.6 and 4.7 comparison covering API migration, token costs, and workload testing.AI Model Comparison

Claude Opus 4.7 vs 4.6: API Migration, Token Costs, and When to Upgrade

Moving an existing Claude Opus 4.6 integration to 4.7 changes thinking controls, sampling parameters, and token counts. Use these request examples and a controlled workload comparison to decide whether the upgrade pays off.

GPT-6 Astra API pricing overview showing context bands, token categories, and the steps for calculating a request's cost.AI Models

GPT-6 Astra API Pricing: Rates, Caching, and Cost Examples

GPT-6 Astra starts at $10 per million ordinary input tokens and $50 per million output tokens. Calculate request costs with the correct cache categories, context band, and processing tier.

Comparison of Kimi K2.6, DeepSeek V4, GPT-5.5 and Claude Opus 4.7 for API workloadsAI Model Comparison

Kimi K2.6 vs DeepSeek V4 vs GPT-5.5 vs Claude Opus 4.7: API Costs and Coding Tests

Compare Kimi K2.6, DeepSeek V4, GPT-5.5 and Claude Opus 4.7 APIs: context, current token prices, worked costs and tests for choosing a coding model.

Gemini image generation stuck loading on a computer beside a phone with a text-only responseGemini Image

Nano Banana Not Working in Gemini? Fix Stuck or Missing Images

If Nano Banana gets stuck loading or Gemini replies without an image, start with the Images tool and the message on screen. A small comparison can help you decide what to try next.

Sora 2 API workflow for pacing requests, diagnosing errors, and saving completed videosAI Video Generation

Sora 2 API Limits: Rate Limits, 429 Errors, and Download Windows

Check the Sora 2 RPM table and your project's effective limits, retry temporary errors without duplicating jobs, and understand the separate one-hour and 24-hour download windows.

Illustrated comparison of cropping, fixed-area repair, and tracking a moving watermark in a downloaded MP4AI Tools

Sora 2 Watermark Remover: How to Clean Up a Downloaded MP4

A practical Sora 2 watermark removal guide for saved MP4s: copyable FFmpeg commands, crop and repair tradeoffs, moving-watermark limits, and audio checks.

Claude Sonnet 5 pricing after the canceled September increaseClaude

Claude Sonnet 5 Price Increase Canceled: Current Rates and Billing Math

Anthropic canceled the September 1 Sonnet 5 API price increase. See the current USD rates, what changes in a September budget, and how to reconcile uncached input, cache writes, cache reads, and output.

FLUX.2 Pro and Flex illustrated as production and adjustable-detail workflowsAI Image Generation

FLUX.2 Flex vs Pro: Which Model Is Worth Using?

Start with FLUX.2 Pro for lower generation costs. Test Flex when text or fine details need more control. Compare current BFL pricing, editing costs, and the acceptance rate that would justify paying more.

Choosing a single RTX 5090 or dual GPUs for Qwen3.8-27B based on precision, context, and concurrency.AI Models

Qwen3.8-27B VRAM Guide: RTX 5090, Dual GPUs, and 262K Context

A 32GB RTX 5090 can run quantized Qwen3.8-27B, but fitting 262,144 tokens depends on the weights, KV cache, runtime, and workload. Use these memory budgets and configuration checks to choose a practical single- or dual-GPU setup.

A request passing through an optional API Management policy and Azure OpenAI deployment checks that can return 429.API

Azure OpenAI TPM Rate Limits: Diagnose 429s Below Quota

Azure OpenAI can return 429 while billed token usage looks low. Compare the receiving deployment’s allocation with response headers, account for estimated tokens and request bursts, and verify recovery without hiding failures behind retries.

GPT-6 Astra request workflow linking the service, project key, and model to response verificationAI API

GPT-6 Astra API Access: Check Your Project and Make a First Request

GPT-6 Astra API access depends on the service, organization, and project behind your key. Use a complete first-request example to check access, read the response correctly, and identify the next step when a call fails.

Astra limits separated into Chat message caps, the shared Work and Codex allowance, and API billing and throughputAI Development Tools

Codex Usage Limits: Check Your Astra Allowance and Get Back to Work

Check your Codex allowance and reset time, then decide whether to wait, use a banked reset, or buy more usage. GPT-6 Astra has different limits in Chat, Work, and the API.

Sora 2 API removal scheduled for September 24, 2026, separate from the April 26 Sora app discontinuationAI Video Generation

Sora 2 API shutdown: the September 24 deadline and what to do now

OpenAI has scheduled the Sora 2 models and Videos API for removal on September 24, 2026. Here is how to separate the app, direct API, and Microsoft Foundry deadlines—and preserve your work before switching providers.

Sora 2 API migration to Veo 3.1 with the scheduled September 24, 2026 removal date.AI Video Generation

Sora 2 vs. Veo 3.1 API: Costs, Compatibility, and Migration

Veo 3.1 can replace short-video generation in a Sora integration, but requests, job tracking, assets, and output constraints change. Compare current API costs and follow the migration through to a saved video.

Illustration for creating and checking transparent PNG assetsAI Image Editing

Make a Transparent PNG with GPT Image 2—or Remove an Existing Background

A reusable cutout needs more than a .png extension. Start with the right creation method, then check the downloaded file and the background where it will actually appear.

Gemini 3.8 Flash and Claude Fable 5.1 API decision visual for developersAI Model Comparison

Gemini 3.8 Flash vs Claude Fable 5.1: Which API Should You Test First?

Gemini 3.8 Flash is the lower-cost first test for most high-volume workloads. Claude Fable 5.1 is a targeted challenger for demanding long-running work—but only your own acceptance, latency, and repair data can justify the premium.

Workload decision board comparing GPT-6 Astra and Claude Fable 5.1 across access, API cost, data rules, and acceptance testingAI Model Comparison

GPT-6 Astra vs Claude Fable 5.1: Choose by Workload

Astra and Fable share the same headline input and output rates, but cache pricing, long-context billing, access, and safeguards can reverse the cheaper or safer choice.

GPT-6 Astra computer-use overview showing integration choices, the Responses API loop, permission checks, state tracking, and outcome verificationAI Development Tools

GPT-6 Astra Computer Use: Build the Right Responses API Loop

GPT-6 Astra can plan multi-step work across browser and desktop interfaces, but your application still owns execution, permission checks, persistent state, and proof that the external task actually succeeded.

Technical illustration for GPT-6 Astra API rate limits, rollout access, and 429/503 recovery.API Guides

OpenAI API Rate Limits: GPT-6 Astra Tiers and 429/503 Recovery

GPT-6 Astra has a public tier table, but your usable capacity still depends on rollout access, organization and project limits, and the response that actually failed.

Claude Free and Pro usage decision based on five-hour and weekly countersClaude

Claude Free vs Pro Usage Limits: Decide With Work, Not Message Counts

Pro gives at least five times Free usage per session during peak periods, but neither plan promises a fixed message count. Use your own interruptions to price the upgrade.

Claude Fable 5.1 and Opus 5 cost decision separating uncached input, cache writes, cache reads, output, and accepted tasksAPI Guides

Claude Fable 5.1 vs Opus 5: Pricing and Cost per Task

Opus 5 costs half as much for uncached input and output. Fable 5.1 costs half as much for cache reads. Your workload and acceptance rate decide which total is lower.

Claude Fable 5.1 plan limits for Pro and Max usersClaude

Claude Fable 5.1 Limits: Pro, Max, and the 50% Cap

Max includes Fable 5.1 within a shared weekly pool, capped at 50% for Fable models. Pro uses paid usage credits from the first Fable request.

Document cards moving through separate Claude chat, Project, context, and usage limit gatesClaude

Does Claude Have Unlimited File Uploads? The Limits That Actually Apply

Claude Projects can hold an unlimited number of files, but that narrow statement comes with per-file and context conditions. Chats, PDFs, images, and account usage each have their own ceilings.

Showing 24 / 439