GitHub deprecated Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code and Claude Opus 4.7 across all Copilot experiences on October 2, 2026, including chat, edits, agent modes and completions. Its suggested replacements are Gemini 3.8 Flash for both Gemini models, Kimi K3 and Claude Opus 5.5, respectively. Copilot Enterprise administrators may need to enable replacement models through policy settings; removing the deprecated models requires no action.
AI21 says it reduced waiting time for high-priority training jobs from up to 72 hours to 12 by adopting Kueue scheduling on its shared GKE GPU fleet. The customer report describes fair admission ordering and topology-aware placement across A3 and A3 Ultra instances using NVIDIA H100 and H200 GPUs. AI21 reports fewer manual interventions and less fragmentation, while saying its total reserved-fleet cost stayed unchanged.
Google's September AI roundup pairs new Gemini models with wider app integrations. The company says Gemini 4 Argon has a million-token context window and is rolling out first to trusted cyber defenders through Fairwind, with broader access planned after feedback on safeguards. Other announcements include Gemini on Windows, more Connected Apps, and WeatherNext 3 forecasts integrated into Search, Gemini, Maps and Cloud.
Web Search API is now available in beta. Web Search API lets your AI agents and applications search the Internet and ground their responses in live information, instead of guessing URLs or relying on a model's training cutoff. At launch, you can choose between three search providers: Ceramic.ai, Exa, and Linkup. All three support Zero Data Retention for requests made through Cloudflare, and all have committed to Cloudflare's verified bot crawling standards. Web Search API runs through AI Gateway, so search requests appear in your gateway logs and are billed to your AI Gateway credits at each provider's list API price, with no additional markup. You can also bring your own provider API key.
Source Cloudflare DevelopersCC BY 4.0 · Publisher excerpt shortened and converted to plain text. Original source license applies.
Google Cloud describes an AI-agent memory design that keeps recent conversations in Memorystore for Valkey and durable facts, preferences and events in AlloyDB. The tutorial combines hybrid retrieval with database-managed embeddings, tenant isolation and memory cleanup. Google reports token and latency reductions from an internal simulated developer workload, while cautioning that savings vary with prompt structure, query frequency and data volume.
Source Google CloudBy Itai Rosenblatt; Paul Ramsey
Vercel says financial AI company Rogo is rebuilding internal applications on its platform and deploying agent-written code in five minutes. The customer case study reports more than 73,000 deployments in one month and six production AI agents handling work including churn analysis and deal support. Vercel also describes automated production-incident triage and remediation using its AI SDK.
Cloudflare has added PiHarness to its Agents SDK, connecting Pi-based agents to Durable Objects. The company says this integration preserves work across interruptions and can keep long-running tasks alive through restarts, crashes and network problems. It supports Workers AI, AI Gateway and existing pi-ai providers. PiHarness remains in beta; Pi Durable is experimental, and the interface is likely to change.
Source Cloudflare DevelopersCC BY 4.0 · Source material adapted into an original NexusTechWire brief. Original source license applies.