Technology intelligence

Clearer
signals.
Brighter
decisions.

NEXUS TECH WIRE

Technology.
Ideas.
People.
A more connected
tomorrow.

AI intelligence

Publisher updates

A useful brief, clear publisher credit, and a direct link to every full article.

Conceptual illustration of connected computing systems and artificial intelligence.
AI-generated illustration Conceptual artwork
AI

AI Gateway - AI Gateway consolidates monthly usage invoice line items and standardizes model names

From the publisher

AI Gateway monthly usage invoices, issued at the beginning of each month for the previous month's usage, now show a single total cost for each model. These invoices no longer break out input and output token quantities and unit prices into separate line items. This change does not apply to invoices for AI Gateway credit purchases. For example, an invoice that previously included these separate line items: anthropic claude-haiku-4-5-20251001 Input Tokens: 40,000 tokens at $0.000001 ($0.04) anthropic claude-haiku-4-5-20251001 Output Tokens: 24,000 tokens at $0.000005 ($0.12) The updated invoice includes one line item: anthropic/claude-haiku-4.5: $0.16. AI Gateway has also standardized model names across invoices and logs. Model variants that previously appeared with provider-specific version suffixes now use a consistent provider/model identifier.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - Z.ai GLM-5.3 now available on Workers AI

From the publisher

@cf/zai-org/glm-5.3 is now available on Workers AI. It is Z.ai's flagship agentic coding model, built for long-running, tool-driven development workflows rather than single-turn chat. GLM-5.3 uses the same base model as GLM-5.2, with every gain coming from post-training. The results are substantial on coding and agentic benchmarks: Z.ai reports ↗︎ a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, and calls GLM-5.3 the most capable open-weights model for coding. On public benchmarks, it scores 88.2 on Terminal Bench 2.1 (up from 81.0), 28.3 on Terminal Bench 3.0 — open-source state of the art, up from 4.6 — 66.9 on DeepSWE (up from 46.2), 78.1 on FrontierSWE (up from 67.5), and 42.5 on SWE-Marathon (up from 19.4).

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - Z.ai GLM-5.3 Flash now available on Workers AI

From the publisher

@cf/zai-org/glm-5.3-flash is now available on Workers AI. It is the first natively multimodal model in the GLM-5 series, built on a Mixture-of-Experts architecture with 320B total parameters and 18B active per token. GLM-5.3 Flash is the first GLM-family model on Workers AI to support multimodal inputs. It outperforms GLM-5.2 across benchmarks and real-world workloads at a lower price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3 Flash requires the Workers Paid plan or prepaid AI Gateway credits. Use GLM-5.3 Flash through the Workers AI binding (env.AI.run()), the REST API, the OpenAI-compatible endpoint, or AI Gateway. For more information, refer to the GLM-5.3 Flash model page and pricing.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Choose OAuth scopes for Wrangler and the Cloudflare API MCP server

From the publisher

Wrangler and the Cloudflare API MCP server now use optional OAuth scopes. During authorization, you can choose which optional scopes to grant instead of approving every scope requested by each client. The consent dialog now includes the option to edit the permissions you grant to Wrangler or the Cloudflare API MCP server: You can then choose which specific permissions to grant: Required scopes remain selected. Choosing fewer optional scopes limits each tool's access to the permissions needed for your workflow. If a command or tool call needs a scope that you declined, reauthorize the client and grant that scope. For more information, refer to wrangler login and Edit optional permissions.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

AI Gateway - Get 50% off GPT-5.6 Sol through AI Gateway

From the publisher

GPT-5.6 Sol is available through AI Gateway, and for a limited time you can use it at 50% off. If you are already using AI Gateway, point to the openai/gpt-5.6-sol model and the discounted pricing applies automatically — no promo code needed. The promotion is available for Unified Billing users only (not Bring Your Own Keys). Load credits onto AI Gateway and start sending requests to openai/gpt-5.6-sol. Discounted pricing during the promotion: Usage Promotional price Standard price Input $2.50 per 1M tokens $5 per 1M tokens Output $15 per 1M tokens $30 per 1M tokens Cache read $0.25 per 1M tokens $0.50 per 1M tokens The promotion runs through September 18, 2026. After that date, GPT-5.6 Sol requests return to standard pricing. For more details, refer to the Unified Billing documentation and the GPT-5.6 Sol model page.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI
Conceptual illustration of connected computing systems and artificial intelligence.
AI-generated illustration Conceptual artwork

Workers AI - Qwen 3.8 27B now available on Workers AI

From the publisher

@cf/qwen/qwen3.8-27b is now available on Workers AI. Qwen 3.8 27B is a 27-billion-parameter instruction-tuned vision language model from Alibaba's Qwen family. It processes images and text together, with reasoning and function calling for agentic workflows. Key capabilities: Vision: Accept image and text inputs and generate text responses. Reasoning: Support thinking mode for complex, step-by-step problem-solving. Function calling: Build agents that invoke tools and APIs across multiple conversation turns. 262,144 token context window: Retain long conversations and multimodal inputs across extended agent sessions. Use Qwen 3.8 27B through the Workers AI binding (env.AI.run()) or the REST API at /ai/run. You can also use AI Gateway with these endpoints. For more information, refer to the Qwen 3.8 27B model page and pricing.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - DeepSeek V4 Flash and Pro now available on Workers AI

From the publisher

@cf/deepseek-ai/deepseek-v4-pro-0813 and @cf/deepseek-ai/deepseek-v4-flash-0731 are now available on Workers AI. DeepSeek V4 Flash and DeepSeek V4 Pro are the first Workers AI models with a full one million (1,048,576) token context window. Use them for long-horizon agentic workflows, large codebases, and multi-step reasoning that exceed the context limits of every other model hosted on the platform. DeepSeek V4 Flash is the faster, lower-cost sibling. This release supersedes the preview version with substantially enhanced agentic capabilities. Key capabilities: Reasoning: Both models support thinking mode for complex, step-by-step problem-solving. Function calling: Build agents that invoke tools and APIs across multiple conversation turns. Long context: Both models support a full 1,048,576 token context window. Both models require the Workers Paid plan or prepaid AI Gateway credits.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

AI Gateway, Workers AI - Workers AI and AI Gateway unify model access and billing

From the publisher

Workers AI and AI Gateway now provide a unified path for accessing models and managing inference traffic. Use the same AI binding and REST API to call models hosted on Workers AI or by supported third-party providers, with AI Gateway providing observability, logging, caching, security, and billing controls. Unified entrypoints and observability The AI binding supports both Workers AI and third-party models through env.AI.run(). The REST API provides shared /ai/ endpoints with Cloudflare authentication across providers. Route a Workers AI request through AI Gateway by specifying a gateway ID. Use default to automatically create a gateway on the first authenticated request, or specify an existing gateway to separate applications and workloads: const response = await env.AI.run( "@cf/zai-org/glm-5.2", { messages: [{ role: "user", content: "What is the capital of France?"

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

AI Gateway - Track AI spend and catch anomalous usage with User Insights

From the publisher

AI Gateway now includes User Insights, a dashboard that gives you two things at once: clear visibility into how much your organization spends on AI, and a security signal that surfaces users whose usage suddenly looks abnormal. It works on the traffic already flowing through your gateway, so there is no additional setup. On the spend side, User Insights shows organization-wide totals for cost, requests, tokens, and adoption, and lets you drill into an individual user to see their spend, top models and providers, cache hit rate, and more. To attribute usage to individual users, add a user identifier with custom metadata or put your gateway behind Cloudflare Access. On the security side, User Insights baselines each user's normal usage from their 95th percentile (p95) session cost over the last 30 days, then flags sessions that exceed both that baseline and an organization-level threshold.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

AI Gateway, Access - Identity-aware controls are now available in AI Gateway

From the publisher

AI Gateway now integrates with Cloudflare Access, giving you two new capabilities: Protect your gateway endpoint. Put your AI Gateway behind Access so you can set policies that control who is allowed to call a specific gateway's endpoint. Identity-aware controls. When traffic reaches AI Gateway through an Access-protected custom domain, AI Gateway can use the authenticated user's Access identity in logs, analytics, routing, and spend controls. With identity-aware controls, you can set spend limits by authenticated user, control which gateways different users can access, filter logs by user, and build policies without passing user IDs from the client application. AI Gateway adds the verified Access user ID to request metadata as cf.user_id. For setup instructions, refer to Cloudflare Access.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers - AI agents can debug Workers with local tracing

From the publisher

wrangler dev and vite dev automatically capture structured OpenTelemetry traces and correlated console logs during local Worker invocations. Debug with AI agents When the tooling detects an AI agent session, it prints a terminal hint pointing to the Local Explorer API at /cdn-cgi/local/explorer/api. The API serves an OpenAPI schema and exposes a read-only observability query endpoint for discovering telemetry, querying traces and logs, and inspecting binding state. The agent can identify the exact failing operation, fix the code, rerun the request, and verify the result. This debug loop requires no deployment or temporary logs. Inspect traces in Local Explorer Humans can inspect the same traces and correlated console logs in the Local Explorer browser UI. Each trace shows spans, timing, attributes, and errors.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI
Conceptual illustration of connected computing systems and artificial intelligence.
AI-generated illustration Conceptual artwork

Agents, Workers - Agent traces for Think, Flue, and AI SDK instrumented by Agents SDK

From the publisher

Agent tracing is now available for applications built with the Agents SDK. Traces show each agent turn alongside model calls, tool runs, approvals, token usage, and Workers runtime operations. Turn on Workers tracing in your Wrangler configuration: { "$schema": "./node_modules/wrangler/config-schema.json", "observability": { "traces": { "enabled": true } } } [observability.traces] enabled = true Think and Flue applications emit agent traces automatically. For direct AI SDK calls, wrap the AI SDK namespace once. wrapAISDK() supports AI SDK v6 and v7.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Preview: @cloudflare/computer agent runtime

From the publisher

We're releasing an early preview of @cloudflare/computer ↗︎, an open-source agent runtime that gives every agent its own computer. The runtime dynamically orchestrates between fast, efficient isolates and full Linux containers, so the agent always runs on the right compute primitive for the task at hand. @cloudflare/computer provides a virtual filesystem backed by SQLite, which you can populate from cloud storage, source control, or any files you choose. Agents can read, write, and edit files, run shell commands, and interact with Git repositories. All operations are gated, audited, and observed.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - Select models now require the Workers Paid plan

From the publisher

We are limiting Workers Free plan access to a few resource-intensive models so we can prioritize capacity for the broader Workers AI user base. This helps everyone get a more reliable inference experience, with fewer 429 and 3040 (Out of Capacity) errors. The following models now require the Workers Paid plan: @cf/moonshotai/kimi-k2.6 @cf/moonshotai/kimi-k2.7-code @cf/zai-org/glm-5.2 On the Workers Free plan, requests to these models now return a 403 HTTP error (internal error 5035) prompting you to upgrade. The Workers Paid plan starts at $5 per month and still includes the 10,000 free Neurons per day allocation, with usage beyond that billed at each model's pricing. Many models remain available on the Workers Free plan, including: @cf/zai-org/glm-4.7-flash @cf/google/gemma-4-26b-a4b-it @cf/nvidia/nemotron-3-120b-a12b For the full list, refer to the Workers AI model catalog.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Cloudflare MCP servers support the new MCP 2026-07-28 Specification

From the publisher

Cloudflare's product-specific MCP servers now support the new MCP 2026-07-28 Specification. Each request runs on a fresh stateless server without an MCP protocol session or protocol-specific Durable Object. The /mcp endpoint also accepts stateless requests from 2025 Streamable HTTP clients. Most clients can reconnect without configuration changes. Use /mcp for new connections. Historical /sse URLs continue to work as aliases for the same Streamable HTTP handler, but they no longer serve the deprecated HTTP+SSE transport. If a client forces SSE transport, change it to Streamable HTTP or automatic transport detection.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Agents SDK adds MCP Specification 2026-07-28 support

From the publisher

Agents SDK v0.20.0 adds client and server support for the MCP 2026-07-28 release candidate ↗︎. Workers can serve tools, prompts, resources, and elicitation without an MCP transport session or Durable Object. Agents can connect to both MCP 2026-07-28 servers and existing legacy servers. Client support The MCP client manager now uses @modelcontextprotocol/client. For each connection, it probes for MCP 2026-07-28 support with server/discover. If the server does not support the stateless protocol, the client continues with the legacy initialize handshake on the same connection. Existing addMcpServer calls do not need a protocol-version setting or separate clients for each protocol generation. For stateless requests, elicitation uses input_required through multi-round-trip requests (MRTR). The legacy path uses the same form and URL handlers for pushed requests.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Agents SDK reduces MCP schema conversion, adds exposure controls for MCP in Think and Code Mode SDK adds direct host APIs

From the publisher

This release reduces repeated MCP schema conversion and adds an opt-out for Think's automatic MCP tool exposure. It also lets non-AI-SDK hosts invoke the durable Code Mode runtime directly. Control direct MCP tool exposure in Think Agents SDK MCP clients now reuse converted input and output schemas while a live connection keeps the same tool catalog. This avoids converting every MCP JSON Schema to Zod again for each model turn. @cloudflare/think also adds includeMcpTools.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI
Conceptual illustration of connected computing systems and artificial intelligence.
AI-generated illustration Conceptual artwork

Operating AI/ML Workloads on Kubernetes: A Headlamp Plugin for Kubeflow

From the publisher

Kubernetes has quietly become the default platform for AI and machine learning. Whether you run notebook servers for data scientists, schedule distributed training jobs, tune hyperparameters, or orchestrate multi-step ML pipelines, those workloads increasingly land on a Kubernetes cluster. Kubeflow is one of the most popular ways to assemble that stack, and it does so the Kubernetes-native way: every capability is exposed as a Custom Resource Definition (CRD). That design is a gift to cluster operators, because it means ML workloads can be observed and managed with the same primitives as everything else in the cluster. But in practice the specialized ML dashboards that ship with these platforms hide the Kubernetes layer underneath. When a notebook is stuck or a training run fails, the operator is often left dropping back to kubectl to find out what actually happened at the Pod level.

Source KubernetesCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Agents, Workers - Agents can respond to MCP elicitation requests

From the publisher

Agents connected to Model Context Protocol (MCP) servers with addMcpServer can now handle elicitation ↗︎ requests. Elicitation lets an MCP server request user input while it handles a tool call. Form mode collects structured, non-sensitive data. URL mode asks for consent before opening an out-of-band flow, such as third-party authorization or payment.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - Plain text output for Markdown Conversion

From the publisher

The Markdown Conversion service now supports a new output conversion option that controls the format of the converted content. Set output.format to text to receive plain text with Markdown syntax removed. The default value is markdown, so existing conversions are unchanged. Use the env.AI binding: await env.AI.toMarkdown( { name: "page.html", blob: new Blob([html]) }, { conversionOptions: { output: { format: "text" }, }, }, ); await env.AI.toMarkdown( { name: "page.html", blob: new Blob([html]) }, { conversionOptions: { output: { format: "text" }, }, }, ); Or call the REST API: curl https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/tomarkdown \ -H 'Authorization: Bearer {API_TOKEN}' \ -F 'files=@index.html' \ -F 'conversionOptions={"output": {"format": "text"}}' When you request text output, the format field of each result is set to text.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI - Moondream 3.1 now available on Workers AI

From the publisher

Partnering with Moondream ↗︎ to bring their latest model @cf/moondream/moondream3.1-9B-A2B to Workers AI. Moondream 3.1 is a fast vision language model built on a mixture-of-experts architecture with 9B total parameters and 2B active, delivering frontier-level visual reasoning while retaining fast, cost-efficient inference. Moondream 3.1 is designed for real-world vision tasks, with a 32K token context window for handling complex queries and structured outputs.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Workers AI, AI Search - Workers AI toMarkdown and AI Search now supports GIF and BMP image conversion

From the publisher

Workers AI Markdown conversion (toMarkdown) now supports .gif and .bmp image files, in addition to the JPEG, PNG, WebP, and SVG formats already supported. GIF and BMP files run through the same image pipeline as other formats. Each image is resized if needed (and for animated GIFs, only the first frame is used), then passed to an object-detection model to identify what it contains. Those detected objects prompt a vision model that writes a natural-language description of the image, which becomes searchable, machine-readable Markdown. AI Search uses toMarkdown automatically to process the files it ingests, so any .gif and .bmp files are included the next time your index syncs, with no configuration changes required. This helps when your content mixes formats, for example a support knowledge base full of screenshots or an archive of BMP scans.

Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
AI

Open source maintainership in the age of AI

From the publisher

AI has really changed the game around software development. More people are leveraging AI than ever to contribute patches to projects they use. To me, this is a good thing as more folks will contribute patches rather than fork or not fix them. The main problem is that AI has made generating code fast but there has been very little improvement in maintaining code bases. In this post, we will highlight the ways the Kubernetes community is adapting to the world of AI assisted coding. The first step of this journey was to develop an AI policy. This seems mundane and bureaucratic but there were many PRs that derailed into discussions around AI usage. The AI policy helps steer the conversation around the project's stance on AI and provides a clear signal to contributors on how to use these tools responsibly.

Source KubernetesCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.

Read full article
Latest collection

Original briefs are AI-assisted and checked against the linked source. Publisher excerpts are labeled separately; each story keeps its original date and article link.