Thursday, September 3, 2026

What are Ollama latest Plans

Ollama’s Pro, Max, and Team plans now use transparent per-token pricing. Based on your feedback, every plan includes a monthly pool of usage credits.

 

very plan includes:

High-performance access to the latest open models, at published per-token rates

Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools

Monthly usage credits included with every plan

Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models

No service fees or hidden limits

New plans

Ollama’s new Pro, Max, and Team plans include a pool of usage credits that refreshes every month. Usage is consumed per token, at the rates published on Ollama’s pricing page and on each model’s page.

 

Pro: $20/month, includes $60 of monthly usage

Max: $100/month, includes $300 of monthly usage

Team: $500/month, includes $1,000 of shared monthly usage for unlimited users

 

The free plan now includes a small amount of monthly usage for a set of starter models. Add usage credits to access all models on the free plan with pay-as-you-go pricing: no service fees, no subscription required.

No hidden fees or limits

Ollama’s new pricing has no service fees and no 5-hour or weekly limits. Each plan’s monthly pool refreshes automatically, and when you use it up, you can keep going at the same per-token rate.

 

Every request runs on dedicated compute in the US and Europe, plus Singapore for a limited set of Qwen models, with zero data retention. We don’t log your prompts, and we never train on your data. You can see exactly what each request cost in your account.

Team plan is now available

Ollama’s Team plan is now available for signup with introductory pricing of $500/month:

 

$1,000 of shared included monthly usage, at published per-token rates

One pool of usage credits shared across your organization with no per-seat limits

Invite unlimited users, and view everyone’s usage in one place

 

Sunday, August 30, 2026

What are the major vLLM configuration parameters for tuning

* **`MAX_MODEL_LEN="32767"`**: Sets the maximum token context window (prompts plus generated answers) the engine will handle. Restricting this to 32,767 tokens instead of the native ultra-large context (like 256K) drastically reduces the GPU memory required by the KV cache, allowing for higher concurrency and faster boot times.

* **`QUANTIZATION_TYPE="fp8"`**: Specifies 8-bit floating-point (FP8) precision for the model weights. This compresses the size of the model down significantly, cutting memory usage roughly in half and accelerating inference speed with minimal loss in model intelligence.

* **`KV_CACHE_DTYPE="fp8"`**: Compresses the Key-Value (KV) cache storage into FP8 format instead of standard FP16/BF16. Because the KV cache grows rapidly with long conversations, this halves its memory footprint, freeing up space to hold more concurrent users and longer contexts.

* **`GPU_MEM_UTIL="0.95"`**: Tells vLLM to claim and lock down **95%** of the total available GPU VRAM on your Cloud Run instance. The remaining 5% is left as breathing room for PyTorch operations and context switching to prevent out-of-memory (OOM) crashes.

* **`TENSOR_PARALLEL_SIZE="1"`**: Determines how many GPUs the model is split across. Set to `1` because your Cloud Run service instance is provisioned with a single GPU, meaning the entire model runs on that single card.

* **`MAX_NUM_SEQS="16"`**: Caps the maximum number of simultaneous sequences (requests) vLLM will batch and process together in a single step [cite: . This prevents memory spikes under heavy web traffic by queuing excess incoming chat requests.



My Cloud Run hangs while downloading Large OpenWeight model Is it normal?

The output is like below 

ravi_retheesh@cloudshell:~ (gemmabigquerymcp)$ gcloud builds submit --project="${GOOGLE_CLOUD_PROJECT}" --region="${GOOGLE_CLOUD_REGION}" --no-source \

substitutions="_MODEL_NAME=${MODEL_NAME},_GCS_MODEL_LOCATION=${GCS_MODEL_LOCATION}" \


    --config=/dev/stdin <<'EOF'


steps:


- name: 'gcr.io/google.com/cloudsdktool/google-cloud-cli:slim'


  entrypoint: 'bash'


  args:


  - '-c'


  - |


    if [[ "$_GCS_MODEL_LOCATION" == *"vertex-model-garden-public-us"* ]]; then


      echo "Using the public cache bucket."


      exit 0


    fi


    gcloud config set storage/parallel_composite_upload_enabled True


    gcloud config set storage/parallel_composite_upload_threshold 150M


    gcloud config set storage/sliced_object_download_threshold 150M


    MODEL_NAME="$_MODEL_NAME"


    SHORT_NAME="$${MODEL_NAME#*/}"


    gcloud storage cp -r -D "gs://vertex-model-garden-public-us/gemma4/$${SHORT_NAME}" "$_GCS_MODEL_LOCATION"


EOF


Created [https://cloudbuild.googleapis.com/v1/projects/gemmabigquerymcp/locations/us-central1/builds/21faa182-824e-42fb-b292-d4368222d847].


Logs are available at [ https://console.cloud.google.com/cloud-build/builds;region=us-central1/21faa182-824e-42fb-b292-d4368222d847?project=593821728960 ].


Waiting for build to complete. Polling interval: 1 second(s).


----------------------------- REMOTE BUILD OUTPUT ------------------------------


starting build "21faa182-824e-42fb-b292-d4368222d847"




FETCHSOURCE


BUILD


Pulling image: gcr.io/google.com/cloudsdktool/google-cloud-cli:slim


slim: Pulling from google.com/cloudsdktool/google-cloud-cli


6310eb16bf42: Pulling fs layer


488f94d8fb5f: Pulling fs layer


4f4fb700ef54: Pulling fs layer


3cbf6ceea8d8: Pulling fs layer


8421ab9bea9a: Pulling fs layer


488f94d8fb5f: Verifying Checksum


488f94d8fb5f: Download complete


8421ab9bea9a: Verifying Checksum


8421ab9bea9a: Download complete


4f4fb700ef54: Verifying Checksum


4f4fb700ef54: Download complete


6310eb16bf42: Verifying Checksum


6310eb16bf42: Download complete


6310eb16bf42: Pull complete


488f94d8fb5f: Pull complete


4f4fb700ef54: Pull complete


3cbf6ceea8d8: Verifying Checksum


3cbf6ceea8d8: Download complete


3cbf6ceea8d8: Pull complete


8421ab9bea9a: Pull complete


Digest: sha256:14026c08cf391812086012d6c5f14078221ce1b46be0df419f7f31ec9f69c712


Status: Downloaded newer image for gcr.io/google.com/cloudsdktool/google-cloud-cli:slim


gcr.io/google.com/cloudsdktool/google-cloud-cli:slim


Updated property [storage/parallel_composite_upload_enabled].


Updated property [storage/parallel_composite_upload_threshold].


Updated property [storage/sliced_object_download_threshold].


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/README.md to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/README.md


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/chat_template.jinja to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/chat_template.jinja


  


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/config.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/config.json


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/generation_config.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/generation_config.json


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/model-00001-of-00002.safetensors to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/model-00001-of-00002.safetensors


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/model-00002-of-00002.safetensors to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/model-00002-of-00002.safetensors


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/model.safetensors.index.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/model.safetensors.index.json


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/processor_config.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/processor_config.json


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/tokenizer.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/tokenizer.json


Copying gs://vertex-model-garden-public-us/gemma4/gemma-4-31B-it/tokenizer_config.json to gs://gemmabigquerymcp-us-central1-hf-model-cache/model-cache/google/gemma-4-31B-it/tokenizer_config.json 







Yes, hanging at this point is completely expected.


What you are looking at in your terminal output is the script successfully copying the smaller configuration files (README.md, tokenizers, json configs), and it has now reached the heavy .safetensors files.


What is happening internally?

Massive File Sizes: The Gemma 4 31B model consists of multi-gigabyte weight tensors split across multiple large files (e.g., model-00001-of-00002.safetensors and model-00002-of-00002.safetensors), each weighing anywhere from 10 GB to 15 GB+ depending on the quantization.


Silent Progress: The standard gcloud storage cp command does not stream a granular, incremental progress bar inside a Cloud Build remote execution log. Because it is processing a massive file payload over the network in total silence, the terminal appears completely frozen.


Is it actually working?

Yes. As long as the build hasn't timed out or thrown an explicit error, the Cloud Build worker is actively pulling the file blocks from Google's public model garden bucket and streaming them into your custom GCS cache bucket (gemmabigquerymcp-us-central1-hf-model-cache).


What should you do?

Wait it out: Depending on network routing inside the us-central1 data center, a multi-gigabyte copy can take anywhere from 2 to 10 minutes.


Verify via Console (Optional): If you want to check if data is actually transferring, you can open a separate Cloud Shell tab and run a check on your destination bucket to see file sizes growing:


Bash

gcloud storage du -s "gs://${MODEL_CACHE_BUCKET}"

Let it run for a few more minutes; it will finish and proceed to the next step once the tensor files are fully copied over.


How to run Large OpenWeight models on Infra like cloud run?

**Yes, this snippet is very good and standard practice** for handling large open-weights models like Gemma 4 on serverless infrastructure like Cloud Run.


### How This Works and Why It's Designed This Way


* **Bypassing Cloud Build Limits:** Cloud Build has strict storage limits and timeouts. Trying to download a massive 30GB+ model inside a standard build step or pushing it straight into a container image will typically break or exceed size thresholds.

* **Leveraging Google's Public Cache (`vertex-model-garden-public-us`):** Google hosts public copies of model weights inside specialized regional GCS buckets. If your Cloud Run service runs in a `us-` region, the snippet intelligently skips downloading entirely (`exit 0`), allowing your app to stream weights straight from Google’s ultra-fast internal network.

* **Parallel Upload Optimization:** If you are using a non-US region or deploying a custom fine-tuned variant, the `gcloud config set storage/...` commands configure parallel composite uploads and sliced downloads. This breaks large model safetensors files into chunks, downloading them concurrently to bypass single-stream bottlenecks.


---


### Your Options to Run Gemma 4 on Google Cloud Run


When deploying Gemma 4 (such as the 31B Dense or 26B MoE variants) to Cloud Run, you have two primary deployment architectures:


| Approach | Architecture | Best For |

| --- | --- | --- |

| **1. vLLM Container + GCS Cache (The Snippet Method)** | Packages vLLM inside a custom container image, pulls weights directly from GCS into GPU memory using high-performance streaming. | High throughput, production-grade OpenAI-compatible API serving, and maximum token generation speed. |

| **2. Ollama on Cloud Run** | Deploys an Ollama runtime container to Cloud Run backed by serverless GPUs, pulling model weights dynamically on start. | Lightweight setups, rapid prototyping, and simpler configuration files. |


Both approaches require provisioning **NVIDIA RTX Pro 6000 (Blackwell)** or similar serverless GPUs on Cloud Run to supply enough VRAM to hold the model weights.

Do Google Provide vLLM service? How do cache the model weights?


 This command creates a dedicated Google Cloud Storage (GCS) bucket to act as a staging cache for your large AI model weights.



Why is this step performed?

Bypassing Hugging Face Rate Limits / Slowdowns: Downloading massive multi-gigabyte open-weights models directly from Hugging Face every time a Cloud Run container spins up is slow and prone to network timeouts or rate limits.


Optimizing Cold Starts with vLLM: When combined with the Run:ai Model Streamer feature built into vLLM, the container can stream model weight tensors concurrently and directly from GCS into GPU memory.

 This drastically minimizes container boot-up and model-loading times.


Direct VPC Egress: Keeping the bucket in the same single region as your Cloud Run service ensures data flows privately and rapidly over Google's internal network rather than traversing the public internet.


Does Google by default provide a vLLM runtime?

No, Google Cloud does not provide a pre-installed, managed vLLM runtime service out of the box (like a native managed-vllm API endpoint).


Instead, Google provides the underlying infrastructure to run vLLM yourself:


Compute Support: Cloud Run supports attaching high-performance GPUs (such as NVIDIA L4 or RTX Pro 6000) and Google Cloud TPUs, alongside serverless scaling.


Bring-Your-Own-Container (BYOC): To use vLLM on Cloud Run, you package vLLM inside a custom container image (or use open-source deep learning containers/Buildpacks) that pulls the model weights from your GCS bucket upon startup and exposes an OpenAI-compatible API endpoint.

Friday, August 28, 2026

What is PodMonitoring and node-exporter in Kubernetes?

A PodMonitoring is a custom Kubernetes resource used by Google Cloud Managed Service for Prometheus to define how metric data from application pods is scraped and ingested.What It DoesTarget Scraping: Finds specific pods within a Kubernetes namespace using label selectors and scrapes their Prometheus-formatted metrics endpoints.Flow Control: Controls how metrics stream into Cloud Monitoring, allowing you to modify scrape intervals, apply time-series filtering, reduce data rates, and use Prometheus relabeling rules.Namespace Scoping: Scrapes targets strictly within the single namespace where the resource is deployed (you must deploy copies across multiple namespaces to scrape globally, or use a ClusterPodMonitoring resource instead)



Node Exporter acts as the data producer, while a monitoring custom resource acts as the data collector.When you want to monitor hardware and OS metrics (like CPU, memory, and disk usage) of your Google Kubernetes Engine (GKE) nodes, you deploy Node Exporter. To actually pull those metrics into Google Cloud, you use a target-scraping custom resource.Because Node Exporter tracks node-level infrastructure rather than isolated application workflows, it connects to Google Cloud Managed Service for Prometheus using two main strategies:1. The Global Connection (ClusterPodMonitoring)While you can use a standard PodMonitoring resource, Google Cloud recommends using a ClusterPodMonitoring resource for Node Exporter.The Problem: Node Exporter runs as a DaemonSet across all nodes, usually isolated in a system namespace (like kube-system or gmp-system). A standard PodMonitoring resource is strictly limited to looking inside its own namespace.The Connection: A ClusterPodMonitoring resource breaks past namespace boundaries. It identifies all Node Exporter pods running across your entire cluster using label selectors (e.g., app.kubernetes.io/name: node-exporter) and scrapes their metric endpoints (typically port 9100).2. The Isolated Connection (PodMonitoring)If you specifically choose to use a standard PodMonitoring resource instead, you must deploy it directly into the exact same namespace where your Node Exporter daemonset lives. It will selectively connect to and scrape only the Node Exporter pods within that specific namespace boundary.What the Configuration Looks LikeTo bridge the gap between Node Exporter and Google Cloud, you apply a manifest that explicitly targets the exporter. Here is the official setup pattern for the Google Cloud Node Exporter integration using a cluster-wide monitoring rule:yamlapiVersion: monitoring.gke.io/v1

kind: ClusterPodMonitoring

metadata:

  name: node-exporter

spec:

  selector:

    matchLabels:

      app.kubernetes.io/name: node-exporter # 1. Finds the Node Exporter pods

  endpoints:

  - port: metrics                         # 2. Targets the exporter's metric port

    interval: 30s                         # 3. Sets how frequently to scrape hardware data

Use code with caution.The Data Pipeline FlowNode Exporter gathers raw kernel, memory, and disk utilization data directly from the host GKE node.Google Cloud's Managed Collector reads your monitoring configuration, discovers the exporter pods via the labels, and scrapes the data.The metrics are forwarded to Google's Monarch database, allowing you to view node health using PromQL in Cloud Monitoring or Grafana.Are you setting up Node Exporter to build a custom cluster dashboard, or are you troubleshooting missing node-level metrics in Cloud Monitoring?


Google Cloud Collect metrics from Exporters using the Managed Service for Prometheus

Main aim is to 

Deploy a GKE instance

Configure the PodMonitoring custom resource and node-exporter tool

Build the GMP binary locally and deploy to the GKE instance

Apply a Prometheus configuration to begin collecting metrics

gcloud auth list

To ingest the metric data emitted by the example application, you use target scraping. Target scraping and metrics ingestion are configured using Kubernetes custom resources. The managed service uses PodMonitoring custom resources (CRs).

A PodMonitoring CR scrapes targets only in the namespace the CR is deployed in. To scrape targets in multiple namespaces, deploy the same PodMonitoring CR in each namespace. You can verify the PodMonitoring resource is installed in the intended namespace by running kubectl get podmonitoring -A.

Google Managed Service for Prometheus (GMP) ingestion of Node Exporter metrics links standard open-source infrastructure monitoring with Google Cloud's fully managed global monitoring backend

Node Exporter: A lightweight agent that runs on your servers or cluster nodes to collect low-level, machine-level hardware and operating system metrics (such as CPU usage, memory, disk I/O, and network traffic). It exposes these metrics locally in a standard Prometheus text format.GMP Binary / Collector: Google's drop-in replacement or forked version of the Prometheus collector/binary. Instead of relying on a locally managed Prometheus time-series database for long-term storage, this specialized binary scrapes the local /metrics endpoint of Node Exporter and forwards (pushes) those metrics directly into Google Cloud's global monitoring database (Monarch / Cloud Monitoring).Viewing the Metrics: Querying the gathered data using standard PromQL inside the Google Cloud Console, Metrics Explorer, or connected dashboards like Grafana







Thursday, August 27, 2026

What is Signoz

 




SigNoz is an open-source observability platform built on OpenTelemetry. We’re building an enterprise-grade alternative to fragmented monitoring stacks, with logs, metrics, traces, alerts, and dashboards in one place.


Choose how to run SigNoz

SigNoz Cloud (Recommended)

Fully managed SigNoz with a 30-day free trial, no credit card required, usage-based pricing that starts at $49, and regional data hosting.


Start free →


Enterprise

Enterprise Cloud, BYOC, or Enterprise Self-Hosted with compliance, support, custom retention, RBAC, ingestion controls, data residency, and region selection.


Explore Enterprise →


Community

Free open-source SigNoz that runs in your own infrastructure. Deploy with Docker, Kubernetes, or Linux and keep full control of your data plane.


Install SigNoz →


What can you monitor?

SigNoz helps teams debug production issues faster by connecting logs, metrics, traces, alerts, dashboards, exceptions, and agent-native workflows in one place.


APM Overview

Monitor service latency, error rate, throughput, Apdex, top endpoints, database calls, and external calls.



Log Management

Ingest, search, aggregate, and correlate logs with traces and metrics using a visual query builder.



Metrics and Dashboards

Build dashboards for application, infrastructure, and custom metrics using Query Builder, PromQL, or ClickHouse SQL.



nfrastructure Monitoring

Monitor Kubernetes clusters, pods, nodes, workloads, and host-level CPU, memory, disk, network, logs, and traces.



LLM and AI Observability

Trace LLM apps, RAG pipelines, prompts, tool calls, tokens, latency, and costs alongside application and infrastructure telemetry.


Agent-Native Observability and MCP

Use the SigNoz MCP server to bring telemetry into coding agents, or use Noz inside SigNoz to investigate incidents, tune alerts, and build dashboards with production context. Noz is available only on SigNoz Cloud.


Distributed Tracing

Follow requests across services with flamegraphs, waterfalls, span events, filters, and trace analytics.


Trace Funnels

Create funnels from traces to understand request-flow drop-offs, failed transitions, and systemic workflow issues.


Learn more: Trace funnels documentation


Also monitor: exceptions, alerts, external APIs, and integrations for OpenTelemetry, Prometheus, Kubernetes, cloud providers, language SDKs, application frameworks, databases, and LLM tools.


Why teams use SigNoz

OpenTelemetry-native

Instrument once with open standards and keep ownership of your telemetry.

Correlated signals

Move from service charts to traces, logs, infra metrics, and exceptions without switching tools.

Single columnar database

Built for high-cardinality, high-volume observability workloads.

Predictable pricing

No per-host pricing, no user-seat pricing, and no special pricing for custom metrics.

Enterprise ready

SOC 2 Type II and HIPAA compliance, RBAC, ingestion controls, custom retention, support, BYOC, and self-hosting.



Getting started

Start on Cloud

Create a managed SigNoz workspace and get your first dashboard without running observability infrastructure.


Start free on SigNoz Cloud


Self-host SigNoz

Run SigNoz in your own infrastructure with Foundry, Docker, Kubernetes, or Linux.


Foundry · Docker · Kubernetes · Linux


Send data

Instrument applications and infrastructure with OpenTelemetry, Prometheus, language SDKs, and integrations.


Instrumentation · Integrations


Comparisons to familiar tools

SigNoz is often adopted by teams moving from a stack of single-purpose tools or commercial platforms with unpredictable pricing.


Prometheus

Good if you just need metrics. SigNoz keeps metrics, logs, traces, dashboards, and alerts together so teams can debug with correlated context.


Jaeger

Jaeger only does distributed tracing. SigNoz adds metrics, logs, trace analytics, dashboards, alerts, exceptions, and trace-to-log workflows.


Elastic

SigNoz uses columnar database for efficient observability analytics and high-cardinality log workloads, with 50% lower resource requirement compared to Elastic during ingestion. Check the detailed study.


Loki

In the linked benchmark, SigNoz indexed all keys in the test setup, while Loki hit max stream errors when more labels were added. Check the detailed study.




What is Continuous Deployment vs Canary vs Blue Green Deployment

Continuous deployment automates code releases to production, while blue-green and canary are release strategies that safely manage how users receive those updates.Continuous DeploymentDefinition: A practice where every code change that passes automated tests is sent directly to production without human approval.Goal: Speed up how fast new features and fixes reach users.Requirement: Needs strong automated testing and quick rollback tools to catch bugs fast.


Blue-Green DeploymentsDefinition: A strategy that uses two identical production environments, usually called Blue (current live version) and Green (new version).How it works: You test the new code in the green environment, then switch all user traffic from blue to green at the same time using a load balancer.Benefits: Near-zero downtime and instant rollback if something breaks by pointing traffic back to blue.Drawback: Costs more because you must run two full copies of your infrastructure.


Canary DeploymentsDefinition: A strategy that releases the new version to a very small subset of users or servers before launching it to everyone.How it works: You route a small slice of traffic (like 5%) to the new build, watch for errors or performance issues, and gradually increase the percentage until everyone is on the new version.Benefits: Lowers risk by limiting the impact of hidden bugs to a small group and saves money by not needing a fully duplicated second environment.Drawback: More complex to route traffic and track metrics accurately.If you'd like, I can:Explain how to use feature flags with these strategiesGive an example of a CI/CD pipeline setup

Wednesday, August 26, 2026

What is AI Sprawl

AI sprawl is the uncontrolled growth of artificial intelligence tools, models, and autonomous agents across a company without central oversight or governance. It happens when different teams adopt AI quickly to solve local problems, leaving IT and security teams without a clear inventory of what is running.You can read more about how this impacts businesses in the IBM Guide on AI Agent Sprawl.How AI Sprawl AppearsShadow AI: Employees use unapproved external AI apps or browser extensions for daily tasks.Embedded AI: Software vendors turn on AI writing, search, or summary features by default in standard workplace apps.Autonomous Agents: Separate teams launch independent AI agents that make decisions, write data, and connect to other software without shared rules.Why It HappensEasy Access: Building or turning on AI features takes seconds and requires no special IT approval.Default Settings: Many business software tools now include AI features automatically.Speed Over Rules: Teams want fast results, so they adopt tools before companies can write clear AI policies.The Main RisksSecurity Gaps: Unmonitored AI integrations create new entry points for data leaks or unauthorized access.Compliance Violations: Autonomous agents may pull private or sensitive user data into unsecured workflows.Conflicting Results: Different tools give different answers or duplicate work, which actually slows teams down.You can explore detection methods through platforms like Reco Security or identity management solutions from Okta.Would you like to know how to audit AI tools in your workplace or how to build an AI governance policy?

Tuesday, August 25, 2026

What are Playwright test agents, Are they for developing tests or running and analyzing them?

 Playwright test agents are LLM-driven AI components (Planner, Generator, and Healer) used primarily for developing and maintaining tests rather than just running them.What Are Playwright Test Agents?Playwright test agents are specialized artificial intelligence tools integrated into Playwright. They interact directly with real browser sessions and live DOM states to automate the test creation and maintenance lifecycle.The three core agents include:Planner: Explores a running application and builds structured Markdown test plans.Generator: Converts those Markdown plans into executable Playwright test script files.Healer: Diagnoses test failures caused by UI or locator updates and automatically repairs the code.Are They Used for Testing or Developing Tests?They are used for developing and maintaining tests, acting as an intelligent layer that writes and fixes the code executed by Playwright's traditional test runner.If you'd like, I can share:How to install and configure these agents using npx playwright init-agentsBest practices for reviewing agent-generated codeLet me know how you want to proceed!


What is Selenium and Playwright , what are the differences

 Selenium Architecture and History

Architecture: Selenium uses the WebDriver API to interact between web browsers and browser drivers. It operates by translating test cases into JSON and sending them to the browsers, which then execute the commands and send an HTTP response back.

History: Selenium has been in continuous operation and development for 18+ years. As a longstanding open source project, it offers broad support for browsers/languages, a wide range of community resources and an ecosystem of support.

Playwright Architecture and History

Architecture: Playwright uses a WebSocket connection rather than the WebDriver API and HTTP. This stays open for the duration of the test, so everything is sent on one connection. This is one reason why Playwright’s execution speeds tend to be faster.

History: Playwright is fairly new to the automation scene. It is faster than Selenium and has capabilities that Selenium lacks, but it does not yet have as broad a range of support for browsers/languages or community support. It is open source and backed by Microsoft.

Comparing Playwright vs Selenium Features

It’s important to consider your own needs and pain points when choosing your next test automation framework. The table below will help you compare Playwright vs Selenium.


Criteria Playwright Selenium

Browser Support Chromium, Firefox, and WebKit (note: Playwright tests browser projects, not stock browsers) Chrome, Safari, Firefox, Opera, Edge, and IE

Language Support Java, Python, .NET C#, TypeScript and JavaScript. Java, Python, C#, Ruby, Perl, PHP, and JavaScript

Test Runner Frameworks Support Jest/Jasmine, AVA, Mocha, and Vitest Jest/Jasmine, Mocha, WebDriver IO, Protractor, TestNG, JUnit, and NUnit

Operating System Support Windows, Mac OS and Linux Windows, Mac OS, Linux and Solaris

Architecture Headless browser with event-driven architecture 4-layer architecture (Selenium Client Library, JSON Wire Protocol, Browser Drivers and Browsers)

Integration with CI Yes Yes

Prerequisites NodeJS Selenium Bindings (for your language), Browser Drivers and Selenium Standalone Server

Real Device Support Native mobile emulation (and experimental real Android support) Real device clouds and remote servers

Community Support Smaller but growing set of community resources Large, established collection of documentation and support options

Open Source Free and open source, backed by Microsoft Free and open source, backed by large community



What is Fixture in PlayWright . How it is different from Hooks

 




A fixture in Playwright is a piece of environment or data that is prepared in advance for a test to use. It is the core concept behind Playwright's built-in dependency injection system, allowing you to isolate and establish a clean, predictable state for every single test case.Instead of writing repetitive code to open a browser, create a page, or configure headers inside traditional beforeEach and afterEach hooks, you simply ask for what your test needs by passing it as an argument.How Fixtures WorkFixtures encapsulate both the setup and teardown logic within a single function using the await use() statement.Code written before await use() acts as the setup phase.The await use() function hands control over to the test itself (and passes any necessary values).Code written after await use() automatically acts as the cleanup/teardown phase once the test finishes.



Example of a Built-In FixtureWhen you write a standard Playwright test, you are already using fixtures, most notably the { page } fixture:javascriptimport { test, expect } from '@playwright/test';


// { page } is a built-in fixture automatically injected into the test

test('basic test', async ({ page }) => {

  await page.goto('https://playwright.dev/');

  await expect(page).toHaveTitle(/Playwright/);

});



Behind the scenes, Playwright automatically launches the browser, creates a unique browser context (like an incognito window), and passes a new page tab into your test block without you needing to manage the browser lifecycle manually.Primary Built-In FixturesPlaywright offers several pre-configured fixtures out of the box:page: An isolated browser tab instance for the specific test run.context: An isolated browser context (session) that handles unique cookies and storage states.browser: The actual browser instance shared across multiple tests in a single worker process.request: A pre-configured API request context instance to handle backend API testing calls.Why Use Fixtures Over Hooks (beforeEach / afterEach)?The Playwright documentation recommends fixtures over traditional hooks because they provide distinct architectural advantages:🏃 Lazy and On-Demand: Fixtures only execute if a specific test explicitly requests them in its arguments. Traditional hooks run before every test regardless of whether the test needs that setup.🧩 Composable: Fixtures can depend on other fixtures. For instance, a custom authenticatedPage fixture can automatically consume and build upon the built-in page fixture.🔀 Scoped Flexibility: They can be configured as test-scoped (recreated for every test case for absolute isolation) or worker-scoped (created once and shared across multiple parallel tests for expensive tasks like setting up a database connection).🛠️ Customizable: You can easily build custom fixtures using test.extend() to encapsulate Page Object Models (POM), log-in sessions, or global mock states into single, reusable arguments.Would you like to see a practical code example of how to create a custom fixture (like an authenticated session or a Page Object Model fixture), or are you trying to migrate existing beforeEach hooks to a fixture-based setup?

Monday, August 24, 2026

Google Cloud Storage - Creating bucket and related

Cloud Storage allows world-wide storage and retrieval of any amount of data at any time. You can use Cloud Storage for a range of scenarios including serving website content, storing data for archival and disaster recovery, or distributing large data objects to users via direct download.



Create a storage bucket

Upload objects to the bucket

Create folders and subfolders in the bucket

Make objects in a storage bucket publicly accessible

I needed to click on the Google Cloud Shell from the top right of the Google Cloud Console. 


Cloud Shell

Manage your infrastructure and develop your applications from any browser with Cloud Shell.

Cloud Shell comes with Cloud SDK gcloud, Cloud Code, an online Code Editor and other utilities pre-installed, fully authenticated and up-to-date. Learn more.

 Cloud Shell is free for all users.

 Gemini Code Assist is enabled in Cloud Shell Editor by default for all users.


Authorize Cloud Shell

Cloud Shell needs permission to use your credentials to make Google Cloud API calls.

Click Authorize to grant permission to this and future calls.


STEP: List the gcloud 

student_01_2cab0553971a@cloudshell:~ (qwiklabs-gcp-01-b6b0a57c3e87)$ gcloud auth list

Credentialed Accounts


ACTIVE: *

ACCOUNT: student-01-2cab0553971a@qwiklabs.net

To set the active account, run:

    $ gcloud config set account `ACCOUNT`


student_01_2cab0553971a@cloudshell:~ (qwiklabs-gcp-01-b6b0a57c3e87)$ 

Create a bucket

In this lab you use gcloud storage commands.

When you create a bucket you must follow the universal bucket naming rules, below.


Bucket naming rules


Do not include sensitive information in the bucket name, because the bucket namespace is global and publicly visible.

Bucket names must contain only lowercase letters, numbers, dashes (-), underscores (_), and dots (.). Names containing dots require verification.

Bucket names must start and end with a number or letter.

Bucket names must contain 3 to 63 characters. Names containing dots can contain up to 222 characters, but each dot-separated component can be no longer than 63 characters.

Bucket names cannot be represented as an IP address in dotted-decimal notation (for example, 192.168.5.4).

Bucket names cannot begin with the "goog" prefix.

Bucket names cannot contain "google" or close misspellings of "google".

Also, for DNS compliance and future compatibility, you should not use underscores (_) or have a period adjacent to another period or dash. For example, ".." or "-." or ".-" are not valid in DNS names.


Use the make bucket (buckets create) command to make a bucket, replacing <YOUR_BUCKET_NAME> with a unique name that follows the bucket naming rules:

gcloud storage buckets create gs://photobucket_rr

gcloud storage cp ada.jpg gs://photobucket_rr

gcloud storage cp -r gs://photobucket_rr/ada.jpg .

gcloud storage cp gs://photobucket_rr/ada.jpg gs://photobucket_rr/image-folder/

gcloud storage ls gs://photobucket_rr

gcloud storage ls -l gs://photobucket_rr/ada.jpg

gcloud storage objects update gs://photobucket_rr/ada.jpg --add-acl-grant=entity=allUsers,role=READER

gcloud storage objects update gs://photobucket_rr/ada.jpg --remove-acl-grant=allUsers

gcloud storage rm gs://photobucket_rr/ada.jpg


Sunday, August 23, 2026

Differences between gRPC, SSE, WebSockets

 The primary difference is their architectural purpose: gRPC is a high-performance framework designed mainly for internal backend microservices, WebSockets provide a persistent channel for two-way (bidirectional) web apps, and SSE (Server-Sent Events) is a simple protocol for one-way (server-to-client) streaming.


gRPC (Google Remote Procedure Call)How it works: A client invokes a function on a remote server as if it were a local function call. It relies strictly on HTTP/2, utilizing features like multiplexing to send multiple requests over one connection without blocking.Payload: Uses Protocol Buffers, which serialize data into a highly compressed binary format instead of readable text. This makes it lightning-fast but harder to debug manually.Strengths: Strict schema enforcement, type safety, and extreme data efficiency for internal backend networks.


WebSocketsHow it works: The client initiates a standard HTTP request and requests a protocol upgrade. Once accepted, the connection morphs into a persistent, raw TCP tunnel where both parties can push messages at any time simultaneously.Payload: Completely flexible. You can stream plain text, raw JSON strings, or binary blobs.Strengths: Low latency for high-frequency, two-way browser data exchanges.


SSE (Server-Sent Events)How it works: The client opens a standard, long-lived HTTP connection using the native browser EventSource interface. The server leaves this response window open indefinitely, pushing text events whenever new data updates arrive.Payload: Text-only. If you need to send binary data, it must be base64 encoded.Strengths: Minimal configuration, standard HTTP firewall compliance, and native handling of drops/reconnections out of the box


Which One to Choose?Choose gRPC if you are linking backend-to-backend infrastructure or mobile apps where network bandwidth and CPU cycles are highly constrained.Choose WebSockets if your client needs to constantly stream data back up to the server while receiving updates, such as inside a fast-paced multiplayer web game or a shared document editor.Choose SSE if your client just needs to sit back and listen to an outgoing feed, such as tracking a live sports score, waiting for system push notifications, or streaming real-time tokens from a Generative AI text API.

Playwright CLI vs standard cli

 Playwright-cli is a terminal-native command-line interface specifically built for AI coding agents to control web browsers. Developed by Microsoft as part of the official Playwright project, it allows terminal-based AI tools to click buttons, take screenshots, navigate pages, and extract data using lightweight shell commands rather than heavy API integrations. [1, 2, 3]  

Why  Was Built 

Before its launch, AI agents used the  Model Context Protocol (MCP)  to automate browsers. However, MCP is highly "token-hungry" because it continuously feeds large tool schemas and verbose webpage details into the AI's limited context window. [3, 4]  

The new  fixes this by introducing Skill-Based Workflows. Instead of sending massive webpage structures back and forth, the agent runs concise terminal commands, saving up to 70–80% on AI token costs. [3, 5, 6]  

Standard CLI () vs. New CLI () 

It is important not to confuse the new agent-focused tool with the traditional developer CLI: [6]  


| Feature | Standard CLI () | New Agent CLI ()  |

| --- | --- | --- |

| Target User | Human developers | AI Coding Agents (e.g., Claude Code, Copilot, Cursor)  |

| Primary Use | Running end-to-end test suites and debugging | Browser exploration and live UI automation  |

| Output Type | Human-readable test reports and code generation UI | Machine-readable YAML snapshots and local files  |

| Token Impact | None | Exceptionally low (saves heavy files to disk instead of LLM context)  |


How It Works (The Core "Skills") 

When you install the CLI, you can generate a  file using the command . This file functions as onboarding documentation that teaches the AI agent exactly what commands it is allowed to run. [3, 7]  

Common terminal actions include: 


• Opening a page:  

• Clicking an element:  

• Capturing the state:  

• Taking a visual check:  [6, 7]  


How to Install It 

The CLI can be installed globally via Node.js package manager: [8]  

Are you trying to configure  to work with a specific AI coding agent (like Claude Code or Cursor), or are you looking for traditional Playwright commands to run your own automated tests? 

AI responses may include mistakes.


[1] https://playwright.dev/agent-cli/introduction

[2] https://playwright-cli.com/

[3] https://www.youtube.com/watch?v=OaFmRHiKp68

[4] https://www.youtube.com/watch?v=CVxEOfGu7Nw

[5] https://playwright.dev/python/docs/getting-started-cli

[6] https://testdino.com/blog/playwright-cli

[7] https://testcollab.com/blog/playwright-cli

[8] https://playwright.dev/docs/getting-started-cli


What are Playwright Agents

 Playwright Agents are Large Language Model (LLM)-driven AI tools embedded natively into the Playwright test automation framework. Released in late 2025 (v1.56), they shift the testing paradigm from manually writing hardcoded test scripts ("how" to test) to describing goals in natural language ("what" to test).Unlike generic code generation tools that predict code based on abstract training data, Playwright Agents interact with real, live browser sessions, inspecting the actual Document Object Model (DOM) and accessibility trees to plan, write, and execute tests.The Three Core Playwright AgentsPlaywright comes with three specialized built-in agents that work independently, sequentially, or together in an autonomous "agentic loop" to handle the full testing lifecycle:🎭 Planner: Explores your live application URL and builds a structured test plan in Markdown format, identifying core user paths and edge cases.🎭 Generator: Reads the Markdown test plans created by the Planner and automatically converts them into fully executable, real Playwright test files (.spec.ts) containing proper selectors and assertions.🎭 Healer: Monitors the execution of the test suite. If a test fails due to a UI change or broken locator, the Healer replays the steps, identifies the change, suggests a patch, and repairs the test autonomously.Architecture and Integration OptionsPlaywright supports two distinct approaches for connecting Large Language Models to web automation:Playwright Model Context Protocol (MCP): Best for specialized agentic loops and exploratory automation. It allows LLMs to persistently inspect page structures, but has a higher token consumption cost due to rich context payloads.Playwright CLI: Designed for coding assistants (like GitHub Copilot or Claude Code). It is highly token-efficient, leveraging concise command-line tools and loaded skills on demand rather than heavy DOM schemas.Key BenefitsSelf-Healing Capabilities: Reduces test suite maintenance by fixing broken CSS/XPath selectors automatically when UI layouts change.Accelerated Bootstrapping: Speeds up development by converting high-level intent into working TypeScript tests in seconds.Real-world Verification: Operates within actual browser environments, executing and validating assertions against live DOM states.If you want to try them out, let me know:Which programming language or framework flavor you are using.Your current code editor (e.g., VS Code).If you want a quick setup guide to configure your first agentic loop.



Thursday, August 20, 2026

Is Playwrite MCP server is safe w.r.to Data?

 Yes, the Playwright Model Context Protocol (MCP) server primarily operates via standard input/output (stdio) and runs completely locally on your machine without connecting to any external cloud-hosted servers.However, because its core purpose is browser automation, the browser instance it controls will connect to external web servers whenever you command it to navigate to a live website.🌐 Understanding How Playwright MCP ConnectsLocal stdio Architecture: When configured in AI applications like Claude Desktop, Cursor, or GitHub Copilot CLI, the server runs entirely as a local sub-process. Communication between your AI app and the Playwright tool occurs over your computer's local stdin and stdout channels. No data from this communication channel is broadcast to the internet.Optional Local Network Modes: The Playwright MCP server can also be configured to run as a local HTTP/SSE server. Even in this mode, it is hosted locally on your device (localhost), though it is accessible to other local IDE tools or custom clients.Browser Internet Traffic: While the MCP connection itself is entirely isolated, the Playwright browser instance (Chromium/Chrome) will connect to external servers whenever the AI instructs it to browse a live URL (e.g., executing a command to scrape data from an external website or test an online application).🛡️ Enterprise Security & Data IsolationBecause it operates locally via stdio, none of your local application context, files, or login tokens are transmitted to an external MCP provider hosting server.However, you should keep the following two data flows in mind:The AI Model Provider: Any text, code, or page data that the Playwright MCP server scrapes from your browser window is fed back into your AI client, which then sends it to your AI provider (like Anthropic or OpenAI) to analyze the page content.Local Application Isolation: If you instruct the AI to browse a local development server (http://localhost:3000), the network traffic remains entirely within your local machine.Would you like assistance configuring the claude_desktop_config.json file to run Playwright MCP locally via stdio, or are you looking to restrict the browser from accessing specific external domains?

Playwright MCP server

 Introduction

The Playwright MCP server provides browser automation capabilities through the Model Context Protocol, enabling LLMs to interact with web pages using structured accessibility snapshots. It works with VS Code, Cursor, Windsurf, Claude Desktop, and any other MCP client — no vision models required.

Prerequisites

Before you begin, make sure you have the following installed:


Node.js 20 or newer

An MCP client: VS Code, Cursor, Windsurf, Claude Code, Claude Desktop, or similar

Getting Started

Installation

Add the Playwright MCP server to your client using the standard configuration:

{

  "mcpServers": {

    "playwright": {

      "command": "npx",

      "args": [

        "@playwright/mcp@latest"

      ]

    }

  }

}


VS Code

Click one of the buttons below to install directly:

Install in VS Code Install in VS Code Insiders

Or install via the VS Code CLI:

code --add-mcp '{"name":"playwright","command":"npx","args":["@playwright/mcp@latest"]}'

Cursor

Install in Cursor

Or go to Cursor Settings → MCP → Add new MCP Server and use command type with npx @playwright/mcp@latest.

Claude Code

claude mcp add playwright npx @playwright/mcp@latest

Claude Desktop

Follow the MCP install guide and use the standard config above.

Other clients

The standard configuration works with most MCP clients, including Windsurf, Cline, Goose, Kiro, Codex, Copilot CLI, and others. Consult your client's MCP documentation for where to place the config.

First interaction

Once the server is connected, ask your AI assistant to interact with a web page:

Navigate to https://demo.playwright.dev/todomvc and add a few todo items.

The assistant will use Playwright MCP tools to open the browser, navigate to the page, and interact with elements — all through structured accessibility snapshots rather than screenshots.

Core Features

Accessibility snapshots

Playwright MCP operates on the page's accessibility tree, not pixels. When a tool runs, it returns a structured snapshot showing the page elements, their roles, and text content. The LLM uses element references from these snapshots to interact with the page:

- heading "todos" [level=1]

- textbox "What needs to be done?" [ref=e5]

- listitem:

  - checkbox "Toggle Todo" [ref=e10]

  - text: "Buy groceries"

The LLM reads this snapshot and uses ref=e5 to type into the textbox or ref=e10 to check the checkbox.

Interacting with pages

Playwright MCP provides tools for all common browser interactions:

Navigation: Open URLs, go back/forward, reload pages.

Clicking and typing: Click elements, type text, fill forms, select dropdowns.

Screenshots: Capture the current page or specific elements for visual verification.

Keyboard and mouse: Press keys, hover, drag and drop.

Dialogs: Accept or dismiss browser dialogs.

Tabs: Create, close, and switch between browser tabs.

Running Playwright code

For complex interactions that go beyond individual tool calls, use the browser_run_code_unsafe tool to execute Playwright scripts directly. This tool runs arbitrary JavaScript in the Playwright server process and is RCE-equivalent — only enable it for trusted MCP clients:


Run this Playwright code to verify the todo count:

async (page) => {

  const count = await page.getByTestId('todo-count').textContent();

  return count;

}


Network monitoring and mocking

Inspect network traffic and mock API responses:


View network requests: List all requests made since page load.

Mock routes: Set up URL pattern matching to return custom responses.

Console messages: Access browser console output for debugging.

Storage state

Save and restore browser state including cookies and localStorage:


Save state: Persist authentication and session data to a file.

Restore state: Load previously saved state into a new session.

Cookie management: List, get, set, and delete individual cookies.

Configuration

Headed mode

By default, Playwright MCP runs the browser in headed mode so you can see what's happening. To run headless:


{

  "mcpServers": {

    "playwright": {

      "command": "npx",

      "args": [

        "@playwright/mcp@latest",

        "--headless"

      ]

    }

  }

}


Browser selection

Choose which browser to use:


{

  "mcpServers": {

    "playwright": {

      "command": "npx",

      "args": [

        "@playwright/mcp@latest",

        "--browser=firefox"

      ]

    }

  }

}


Supported values: chrome, firefox, webkit, msedge.


User profile

Playwright MCP supports three profile modes:


Persistent (default): Login state and cookies are preserved between sessions. The profile is stored in ms-playwright/mcp-{channel}-{workspace-hash} in your platform's cache directory, so different projects get separate profiles automatically. Override with --user-data-dir.

Isolated: Each session starts fresh. Pass --isolated to enable. You can load initial state with --storage-state.

Browser extension: Connect to your existing browser tabs with the Playwright Extension. Pass --extension to enable.

Configuration file

For advanced configuration, use a JSON config file:


npx @playwright/mcp@latest --config path/to/config.json


The config file supports browser options, context options, network rules, timeouts, and more. See the Playwright MCP repository for the full schema.


Standalone server

When running a headed browser on a system without a display or from IDE worker processes, start the MCP server separately with HTTP transport:


npx @playwright/mcp@latest --port 8931


HTTP sessions use a five-second heartbeat timeout. If your MCP client or proxy does not answer server-initiated pings, set PLAYWRIGHT_MCP_PING_TIMEOUT_MS to a longer timeout in milliseconds. Set it to 0 to disable the heartbeat.


Then point your MCP client to the HTTP endpoint:


{

  "mcpServers": {

    "playwright": {

      "url": "http://localhost:8931/mcp"

    }

  }

}


Quick Reference

Action How to do it

Install server Add standard config to your MCP client

Navigate to a page Ask: "Go to https://example.com"

Click an element Ask: "Click the Submit button"

Fill a form Ask: "Fill in the email field with test@example.com"

Take a screenshot Ask: "Take a screenshot of the page"

Run Playwright code Ask: "Run this Playwright code: ..."

Mock an API Ask: "Mock the /api/users endpoint to return ..."

Use headed mode Default. Pass --headless to disable

Choose a browser Pass --browser=firefox in args


Tuesday, August 18, 2026

How to map Claude extension features to goals ?

 


This is best overview of how to match the goal and the features that Claude provides.