Vercel's September 9 CLI update brings release announcements into the terminal for developers and coding agents. The changelog command returns five recent entries with their complete Markdown by default, with options to change the limit, search by keyword or produce JSON for scripts. The feature requires Vercel CLI 59.6.0 or later.
AI Gateway custom costs now support cache-read and cache-write token rates. This lets custom cost metrics reflect negotiated cache pricing across providers. Add per_cache_read_token or per_cache_write_token to the cf-aig-custom-cost header: { "per_token_in": 0.000001, "per_token_out": 0.000002, "per_cache_read_token": 0.0000001, "per_cache_write_token": 0.0000005 } Cache-token pricing activates when either cache rate is present. An omitted cache rate defaults to per_token_in. If both cache rates are omitted, AI Gateway preserves the existing input and output calculation. Providers can include cache tokens within input tokens or report them separately. AI Gateway automatically accounts for these differences and prevents double-counting. For more information, refer to Custom costs.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Vercel added DeepSeek V4.1 Flash to AI Gateway with combined text and image input, including uses such as interpreting screenshots and charts. The announcement lists a one-million-token context window, responses up to 384,000 tokens, reasoning, tool use and prompt caching. Developers can select the model through Gateway or configure supported coding agents with the latest Vercel CLI.
Vercel added in-chat provider connections to v0 on September 9, beginning with Resend, Amazon OpenSearch, MongoDB Atlas, Algolia and Clerk. When a prompt needs one of these services, v0 offers a connection card and then configures the required environment variables. Vercel says it also loads a provider's published agent skills automatically, when available, to guide generated code.
Vercel's September 8 release note says its deployment step became about 10% faster after it replaced separate routing metadata uploads for each function path with one combined manifest. The company reports an average saving of one second, with large applications saving up to 12 seconds. The optimization applies automatically to all builds and requires no project changes.
Source VercelBy Ali Smesseim; Tim Caswell; Steven Salat; Javi Velasco
Vercel reported on September 8 that Sandbox domain lookups fell from a median 62 milliseconds to 3.4 milliseconds after routing moved from one central store to nearby regional replicas. The change affects public domains created through sandbox.domain() and is applied automatically. Vercel says pricing is unchanged; the reported 18-fold improvement measures domain lookup latency, rather than an entire application's response time.
Source VercelBy Marc Codina Segura; Tom Lienard; Luke Phillips-Sheard; Andy Waller
AI/ML and complex batch workloads continue to push the boundaries of Kubernetes scheduling. Following the foundational workload-centric enhancements introduced in previous releases, Kubernetes v1.37 delivers the next major milestone in the Workload-Aware Scheduling (WAS) journey. In this release, the core Workload and PodGroup APIs—enabling gang scheduling—along with Workload-Aware Preemption (WAP) and shared DRA ResourceClaims for PodGroups, all graduate to Beta, solidifying their role in the Kubernetes ecosystem. To address the hierarchical scheduling requirements of modern high-performance distributed workloads, v1.37 introduces the new CompositePodGroup API. This new API allows expressing multi-level topology constraints, gang scheduling, and preemption policies for complex, heterogeneous groups of Pods.
Source KubernetesCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Google DeepMind introduced AlphaGenome Atlas on September 8 with precomputed predictions for nine billion possible single-letter human DNA changes. A free portal makes the dataset available for academic research. Its AVI score combines AlphaGenome and AlphaMissense predictions to help researchers prioritize variants and inspect their predicted molecular effects. These are model-based research predictions; the announcement describes targeted experiments validating selected findings.
Mistral says it raised €3 billion in a Series D round at a valuation above €21 billion after the investment. Samsung Electronics led the round, with Scaleup Europe Fund and PSG Equity as co-leads. The company plans to expand AI research, computing infrastructure and international commercial operations, supporting its strategy of open-weight models and enterprise systems that customers can control.
Vercel introduced Flat Rate CDN for Pro teams on September 8, replacing usage billing for covered CDN resources with a monthly capacity tier shared across projects. Vercel says temporary spikes will not raise that period's bill, but sustained growth can change the next month's tier. New Pro teams receive it by default; existing teams can opt in or retain usage-based billing.
Vercel made Flat Rate CDN generally available for Pro teams on September 8. The included tier provides one million requests and one terabyte of transfer, shared across a team's projects; larger tiers cost extra. Vercel says temporary excess traffic continues without additional billing under its fair-use rules. Coverage includes CDN requests, specified transfers and related observability events, rather than every platform charge.
Source VercelBy Kacie Fijalkovich; Jeff Pope; Lakshay Bhushan; Sudais Moorad; Kostyantyn Voytenko; Caleb Boyd; Jas Garcha
Miniflare v5 prepares Cloudflare local development tooling for the upcoming cf CLI. Miniflare powers local Workers development behind wrangler dev, the Cloudflare Vite plugin, and @cloudflare/vitest-plugin. Most projects should use those tools instead of depending on Miniflare directly, and Miniflare v5 will not require any action. The most significant change is a new configuration shape which aligns Miniflare with cloudflare.config.ts, the programmatic Cloudflare configuration format now available for testing. Other breaking changes include: Removed deprecated APIs and options, such as legacy alpha D1 bindings. Removed now-unused, internal APIs like wrappedBindings Removed Miniflare's built-in module discovery; higher-level tools like Wrangler and the Vite plugin should be providing the module graph. Moved local-only /cdn-cgi routes under /cdn-cgi/local.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Vercel added OpenAI's GPT Image 2.5 Flare and Sunburst to AI Gateway for image generation and editing from text prompts or reference images. Its announcement recommends Flare for faster generation and Sunburst when editing precision matters more. Vercel describes support for transparent backgrounds, complex layouts and targeted changes that preserve other image content, with both models available in its playground.
Sep 4, 22:26 UTC Resolved - On September 4, 2026, between 20:04 and 22:26 UTC, GitHub Copilot code review experienced an increased failure rate. Affected pull request reviews failed to complete or post review comments.
Sep 4, 22:23 UTC Resolved - On September 4, 2026, between approximately 21:45 and 22:07 UTC, some users experienced errors and elevated latency for repository operations. The incident was fully resolved at 22:23 UTC.
Kubernetes v1.37 promotes the KubeletInUserNamespace feature gate to beta. With this feature enabled, all of the node components (kubelet, CRI and OCI runtimes, CNI plugins, and kube-proxy) can run as a non-root user on the host, using a Linux user namespace. This technique is also known as rootless mode. The work started as an experiment in 2018, and was merged into Kubernetes v1.22 (2021) as an alpha feature (Kubernetes Enhancement Proposal KEP-2033). This feature should not be confused with user namespaces for pods (hostUsers: false with the UserNamespacesSupport feature gate, GA since v1.36), which puts pods in user namespaces but still runs the node components as root. These two features do not conflict. Moreover, they can be combined to nest Kubernetes inside Kubernetes without resorting to the full privileged: true. Why run the node components in a user namespace?
Source KubernetesCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
You can now deploy Workers with larger dependencies, heavier frameworks, and more code without hitting size limits. When you deploy a Worker, Wrangler bundles your code and compresses it before uploading. Previously, Cloudflare checked that compressed size and rejected deploys over 3 MB (Free) or 10 MB (Paid). That limit has been removed. Cloudflare now only checks the uncompressed size of your bundle, which is 64 MiB across all plans. To check your Worker's bundle size before deploying: wrangler deploy --outdir bundled/ --dry-run Total Upload: 259.61 KiB / gzip: 47.23 KiB The Total Upload value is your uncompressed bundle size. This is what counts against the 64 MiB limit. The gzip value is shown for reference but is no longer a limit. For more information, refer to the Worker size limits documentation.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Vercel announced free access to inclusionAI's Ling 3.0 Flash Sante through October 4. The healthcare-focused model is presented for reasoning, research and evidence retrieval, with a 256K-token context and function calling. After the offer ends, the standard model endpoint continues at normal rates, while the endpoint ending in '-free' stops serving requests; free usage remains visible in traces.
You can now manage Email Routing rules that route emails to Workers from your Wrangler configuration. Add literal addresses or a catch-all address to the top-level addresses field: { "$schema": "./node_modules/wrangler/config-schema.json", "name": "invoice-handler", "main": "src/index.ts", // Set this to today's date "compatibility_date": "2026-10-02", "addresses": [ "invoice@yourdomain.com" ] } name = "invoice-handler" main = "src/index.ts" # Set this to today's date compatibility_date = "2026-10-02" addresses = ["invoice@yourdomain.com"] When you run wrangler deploy, Wrangler creates rules for new addresses, updates existing rules managed by the Worker, and removes managed rules that are no longer in the configuration. Wrangler shows the planned changes and asks for confirmation before applying potentially destructive changes.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Vercel's September 4 changelog added OpenAI's GPT 6 Astra to AI Gateway under the model identifier openai/gpt-6-astra. Vercel describes it as designed for extended agent tasks such as software navigation, data analysis and website development. The announcement provides an AI SDK example and coding-agent setup instructions, including links for Codex and Hermes, giving developers several ways to connect the model.
Source VercelBy Zachary Chen; Kevin Dawkins; Jerilyn Zheng
Enterprise customers can now configure a zone's CDN Maximum Upload Size up to 5 GB directly from the Network page in the Cloudflare dashboard. This removes the need to contact your account team or Cloudflare Support when applications need to accept request bodies larger than 500 MB and no greater than 5 GB. The default maximum upload size remains 500 MB. Upload limits above 5 GB still require additional configuration through your account team or Cloudflare Support. Very large uploads may reach connection or read timeouts before reaching the configured size limit. Make sure clients and origins allow enough time to complete the transfer when increasing this setting. Refer to Cache upload limits and Workers request body size limits for details.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Kubernetes 1.37 is here and Dynamic Resource Allocation (DRA) keeps pushing past where it started! This release brings DRA Extended Resource support to GA, a milestone the team has been building toward for three straight releases. Several more features graduate to Beta or GA. A fresh batch of alpha features rounds out the release. I'll dive into what's new for DRA in Kubernetes 1.37! What's stable in 1.37 DRA Extended Resource support has graduated to GA. This is the mechanism that lets DRA drivers satisfy requests made through the traditional extended resource API, think example.com/gpu in a Pod spec, without requiring a separate device plugin alongside the DRA driver. An extended resource name can be set directly on a DeviceClass, and Pods requesting it get matched to a device through DRA with no ResourceClaim needed on the workload's part.
Source KubernetesCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
GitHub reports that a resolved September 3, 2026 incident disrupted Grok models in Copilot, including Grok 4.5 and 4.6, from 13:22 to 17:11 UTC. Requests to those models produced more errors, while other models remained unaffected. GitHub attributed the disruption to its upstream model provider, showed warnings inside affected products, and said service recovered after the provider applied a mitigation.
Google DeepMind and Google Research have introduced WeatherNext 3, which Google says uses live satellite observations to refresh global forecasts hourly. Temperature and moisture forecasts reach a five-kilometer grid, while other variables use coarser resolutions. Google says deployment begins across products including Search, Gemini and Maps, with data also available through cloud services. The announcement directs users to meteorological agencies for official warnings and public-safety advice.
Original briefs are AI-assisted and checked against the linked source. Publisher excerpts are labeled separately; each story keeps its original date and article link.