← Back to live feed

Wednesday, Sep 23, 2026

1
vLLM Adds DiffusionGemma-Jev to Ease Cloud AI DeploymentGOOGL
β–Ά
topics πŸ€– AI tags AIAI InfraAI Inference GOOGL keywords Google

Google now offers a one-command deployment path for the DiffusionGemma-Jev model on Cloud Run, allowing developers to create API-compatible endpoints without providing their own GPUs. This environment provides single step latency between 35 and 60 ms and a throughput of 100 to 123 requests per second for batch sizes of 32. The model is designed to answer yes/no and multiple-choice questions by reading probability distributions in a single denoising step.

The vLLM project announced on September 23 that the engine now supports DiffusionGemma-Jev, enabling the model to return confidence scores by reading from noisy answer slots. Hosting through Google Cloud Run costs approximately $3 per hour and drops to $0 when the system is idle.

Image via @googlegemma on X
Continues from Tuesday, Sep 22
Google Launches One Command DiffusionGemma-Jev Deployment on Cloud Run
2 tweets β€’ 2 sources
See all 4 tweets β†’