The reference_agent is primarily an AI-powered OKF documentation/enrichment pipeline. It does not itself create the Acme Retail metrics/, computations/, policies/, skills/, and attesters/ system.
Its implemented source pipeline is currently BigQuery → LLM → OKF documents, optionally followed by a web-crawling/enrichment pass → OKF documents → regenerated indexes.
BigQuery Dataset
│
▼
┌──────────────────┐
│ BigQuerySource │
└────────┬─────────┘
│
discovers concepts
│
▼
┌──────────────────┐
│ Reference Agent │
│ Gemini LLM │
└────────┬─────────┘
│
┌────────────┼────────────┐
│ │ │
▼ ▼ ▼
metadata existing other
from BQ OKF concepts
│ │ │
└────────────┼────────────┘
▼
OKF Markdown
│
▼
index.md files
│
▼
OKF Bundle
│
optional Web Pass
│
▼
enriched OKF
Google's README explicitly describes the reference agent as a two-pass system:
BQ pass — generates one OKF document per source concept using BigQuery metadata.
Web pass — optionally crawls documentation starting from supplied seed URLs and enriches existing concepts or creates reference documents.
3. The first major component: BigQuerySource
This is where the process starts.
The CLI currently supports:
--source bq
and requires:
--dataset project.dataset
The CLI constructs:
BigQuerySource(
dataset=args.dataset,
billing_project=args.billing_project
)
The BigQuerySource uses the Google BigQuery Python client.
4. What does BigQuerySource discover?
This is important.
It doesn't simply dump the BigQuery schema.
It creates concepts.
For example:
BigQuery Dataset
│
├── BigQuery Table
├── BigQuery Table
├── BigQuery Table
└── BigQuery Table
The source creates a concept such as:
datasets/my_dataset
with:
type = BigQuery Dataset
and a canonical BigQuery resource URI.
For tables it creates concepts like:
tables/orders
tables/customers
tables/products
with:
type = BigQuery Table
5. It even handles sharded BigQuery tables
There is a nice detail in the implementation.
Suppose your dataset has:
events_20260101
events_20260102
events_20260103
...
The source recognizes these as a sharded table family rather than blindly treating every table as an independent concept.
It detects suffixes using a regex and can create a wildcard concept such as:
tables/events_
with a resource like:
.../tables/events_*
and metadata indicating that it is a wildcard/family.
That's useful because an OKF bundle shouldn't necessarily contain thousands of nearly identical documents for date-sharded tables.
6. Then the agent gets raw metadata
The LLM has a tool called:
read_concept_raw()
This retrieves structured metadata for a concept.
For a BigQuery table, that includes things such as:
schema
nested RECORD fields
partitioning
clustering
row counts
timestamps
The tool implementation explicitly describes these capabilities.
So the LLM isn't guessing the schema.
It receives actual source metadata.
7. The LLM itself is a Google ADK Agent
This is a critical architectural point.
The reference agent is not just a Python script calling Gemini once.
It uses:
from google.adk import Agent
and creates an ADK agent.
The default model is:
gemini-flash-latest
The BQ agent is:
okf_bq_reference_agent
with these tools:
list_concepts
read_concept_raw
sample_rows
read_existing_doc
write_concept_doc
So conceptually:
Gemini
│
┌───────────┼───────────┐
▼ ▼ ▼
list_concepts read_raw sample_rows
│ │ │
└───────────┼───────────┘
│
▼
reason about data
│
▼
write_concept_doc
8. What instructions does Gemini receive?
This is where prompts/reference_instruction.md becomes very important.
The agent is explicitly instructed to follow a workflow.
For each concept:
Step 1 — Check whether an OKF document already exists
read_existing_doc(concept_id)
If it exists, the agent is supposed to refine it rather than blindly replace it.
That's a significant design choice.
Step 2 — Read the source metadata
read_concept_raw(concept_id)
This gives the LLM the actual structured information from BigQuery.
Step 3 — Optionally sample data
sample_rows(concept_id, n=3)
The agent is allowed to do this when metadata alone isn't enough.
This is particularly useful when the schema is something like:
customer_id
product_id
status
amount
created_at
The schema tells you what the fields are, but sample rows can help the model understand what they actually represent.
9. Then it looks at other concepts
The LLM calls:
list_concepts()
This gives it the other concepts available in the bundle.
Why?
Because the generated documentation can create cross-links.
For example:
orders
│
├──────────────► customers
│
└──────────────► products
Rather than creating isolated documents:
orders.md
customers.md
products.md
the agent can create relationships between them.
The prompt specifically tells the agent to use the list of concepts to weave cross-links into its prose.
10. Then Gemini writes the OKF document
The final tool is:
write_concept_doc()
The LLM supplies:
concept_id
frontmatter
body
The tool writes the Markdown file.
The OKF document contains YAML frontmatter and Markdown body. The implementation validates that at least type exists and automatically handles metadata such as generated.
For example, the generated document can conceptually look like:
---
type: BigQuery Table
title: Orders
description: Customer orders placed through the retail platform.
resource: https://bigquery.googleapis.com/...
tags:
- orders
- sales
- transactions
generated:
by: reference_agent/gemini-...
at: ...
sources:
- id: bigquery
resource: ...
title: BigQuery metadata
---
The Orders table contains one row per customer order...
# Schema
| Field | Type | Description |
|---|---|---|
| order_id | STRING | Unique order identifier |
| customer_id | STRING | Customer identifier |
| amount | NUMERIC | Order amount |
| created_at | TIMESTAMP | Order creation time |
# Common query patterns
```sql
SELECT ...
The exact content is generated by the model, but the prompt dictates this general structure: description, `# Schema`, and `# Common query patterns`. :contentReference[oaicite:12]{index=12}
---
# 11. So what exactly is AI-generated?
This is worth emphasizing.
The BigQuery API provides:
```text
facts
─────
table name
columns
types
modes
descriptions
partitioning
clustering
row counts
timestamps
The LLM turns those facts into:
knowledge
─────────
human-readable description
semantic interpretation
field explanations
relationships
common query patterns
cross-links
tags
So:
BigQuery Metadata
│
▼
deterministic
metadata
│
▼
Gemini
│
▼
semantic OKF
documentation
This is essentially metadata → semantic knowledge transformation.
12. Then comes the Web Pass
This is the second major capability.
You can supply:
--web-seed https://...
or:
--web-seed-file seeds.txt
The CLI also supports:
--web-max-pages
--web-max-depth
--web-allowed-host
--web-allowed-path-prefix
--web-denied-path-substring
So you might give it:
https://developers.google.com/analytics
as a seed.
The web agent doesn't simply download the page.
It crawls selectively.
13. The web agent decides what to follow
The web ingestion prompt tells it to start with the seed URL and follow links that look useful for the existing concepts.
For example:
Seed
│
▼
GA4 documentation
│
├── schema
│
├── dimensions
│
├── metrics
│
├── query examples
│
└── unrelated pages
The agent might decide:
schema → follow
metrics → follow
query examples → follow
pricing → skip
login → skip
marketing → skip
The crawler has hard limits on:
number of pages
hop depth
allowed hosts
path prefixes
denied path substrings
These are explicitly passed into the web agent.
14. What does the Web Agent do with a page?
For every fetched page, it decides among three possibilities:
A. Enrich an existing concept
For example:
tables/events.md
already exists.
The documentation says:
event_name represents the GA4 event name.
The web agent can augment the existing OKF document with that information.
B. Create a new reference document
For example:
references/event_parameters.md
if the page contains useful material that doesn't naturally belong to one existing concept.
C. Skip it
If it is:
marketing
navigation
login
irrelevant
it can ignore it.
This behavior is explicitly described in the web ingestion workflow.
15. This gives you an important two-pass architecture
The whole thing is:
SOURCE PASS
│
▼
BigQuery API
│
▼
discover concepts
│
▼
raw metadata tools
│
▼
Gemini
│
▼
OKF documents
│
│
▼
WEB ENRICHMENT
│
▼
seed URLs
│
▼
Gemini crawler
│
┌─────────┼─────────┐
▼ ▼ ▼
enrich create skip
concept reference
│ │
└─────────┼─────────┘
▼
OKF Bundle
│
▼
regenerate indexes
The runner implements exactly this sequence: enrich concepts, optionally run the web pass, then regenerate index.md files.
16. What runner.py is doing
runner.py is effectively the orchestrator.
It creates:
BQ ADK Agent
and, if web seeds are supplied:
Web ADK Agent
It then creates ADK Runner instances and sessions for them.
The BQ flow is approximately:
concepts = source.list_concepts()
for concept in concepts:
enrich_concept(concept)
run_web_pass()
regenerate_indexes()
That's essentially what enrich_all() does.
17. One particularly useful feature: enrich only one concept
The CLI supports:
--concept tables/orders
and it is repeatable.
So instead of rebuilding everything:
reference-agent enrich \
--source bq \
--dataset myproject.sales \
--out ./bundle \
--concept tables/orders
you can target a specific concept.
The runner filters the discovered concepts and raises an error if the requested concept doesn't exist.
This is useful for incremental knowledge maintenance.
18. What it does NOT do
This is probably the most important part given your previous question about the Acme Retail bundle.
The reference agent does not appear to automatically discover and create:
metrics/
computations/
policies/
skills/
attesters/
from a BigQuery dataset.
Its documented BQ workflow creates/enriches concepts such as:
datasets/
tables/
and the web pass can create:
references/
The Acme-specific operational artifacts are therefore not something you should assume the reference agent generates automatically.
This is consistent with the current repository architecture and README description of the reference agent.
19. Then what is Acme Retail doing?
This is where the distinction becomes very interesting.
You effectively have:
Reference Agent
DATA DISCOVERY
│
▼
BigQuery metadata
│
▼
semantic enrichment
│
▼
OKF docs
versus the Acme bundle:
GOVERNED BUSINESS KNOWLEDGE
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
Metrics Policies Computations
│ │ │
└─────────────────┼─────────────────┘
▼
Skills
│
▼
Attesters
The latter is an application of the OKF format, not the entire function of reference_agent.
And the current OKF specification is intentionally minimal: Markdown + YAML frontmatter, with no central schema registry or required tooling.
20. The visualization is a separate function
Another important distinction:
reference_agent enrich
and
reference_agent visualize
are separate CLI operations.
The visualization command:
reference-agent visualize \
--bundle ./my_bundle
doesn't involve Gemini.
It reads the OKF bundle and generates a self-contained:
viz.html
The CLI explicitly describes it as:
Generate a self-contained HTML graph view of an OKF bundle.
and reports the number of concepts and edges generated.
So:
reference_agent
│
┌────────────┴────────────┐
│ │
enrich visualize
│ │
▼ ▼
Gemini + BQ OKF bundle
+ Web │
│ ▼
▼ viz.html
OKF files
21. The most important conceptual takeaway
If you're looking at this from your RAG / Agent / Knowledge Graph perspective, I would describe reference_agent as:
An AI-powered knowledge-catalog generation pipeline that converts source-system metadata and authoritative web documentation into interconnected, provenance-aware OKF documents.
It is not primarily a RAG engine.
It is not a vector database.
It is not GraphRAG.
It is not an agent answering end-user questions.
It is essentially a knowledge preparation / knowledge engineering agent.
Raw Enterprise Knowledge
│
┌────────────┴────────────┐
│ │
Structured source Web docs
(currently BQ) │
│ │
▼ ▼
Source tools Web crawler
│ │
└────────────┬────────────┘
▼
Gemini
│
semantic enrichment
│
▼
┌──────────────┐
│ OKF Bundle │
└──────┬───────┘
│
┌──────────┼──────────┐
▼ ▼ ▼
RAG Agent Graph UI
And this last architecture is where I think OKF becomes particularly relevant to the work you've been doing with RAG/GraphRAG/agents: instead of asking your RAG pipeline to discover enterprise semantics every time a user asks a question, you can have a separate knowledge-engineering pipeline create and maintain the semantic layer first.
Then your runtime agent consumes that layer.