Wednesday, September 10, 2025

What is FalkorDB?


• FalkorDB is a graph database built on top of Redis, designed for real-time AI/ML applications.

• It is a fork of RedisGraph (after Redis stopped maintaining RedisGraph in 2023).

• Uses GraphBLAS (linear algebra-based graph processing) for speed.

• Query language: Cypher-like syntax (similar to Neo4j).


Think of it as: Redis (fast in-memory DB) + Graph structure support + AI-friendly features.


⸻


🔹 Advantages of FalkorDB

1. Performance (In-Memory + GraphBLAS)

• Extremely fast queries, thanks to in-memory Redis + linear algebra ops.

• Good for low-latency use cases (e.g., recommendations, fraud detection).

2. Real-time AI/ML Support

• Supports hybrid search (vector embeddings + graph search).

• Can combine semantic search (vector DB) with graph traversals.

3. Cypher Query Language Support

• Developers familiar with Neo4j/Cypher can adapt quickly.

4. Scalability

• Inherits Redis cluster scalability.

• Works well in distributed, high-throughput environments.

5. Open Source & Actively Maintained

• Unlike RedisGraph (which is discontinued), FalkorDB is actively updated.

6. Integration with AI frameworks

• Works nicely with LLMs, recommendation engines, and knowledge graphs.


⸻


🔹 Disadvantages of FalkorDB

1. Memory Intensive

• Like Redis, it stores data in memory (RAM).

• Expensive for very large graphs unless persistence layers are optimized.

2. Younger Ecosystem

• Compared to Neo4j or ArangoDB, community and ecosystem are smaller.

• Fewer third-party integrations, tutorials, and production deployments.

3. Feature Gap vs Neo4j

• Neo4j still has richer tooling (Bloom visualization, enterprise features, plugins).

• FalkorDB is more lightweight.

4. Operational Complexity

• Needs careful memory management and persistence tuning.

• Scaling beyond RAM can be tricky compared to disk-based graph DBs.

5. Limited Query Language Extensions

• Cypher support is partial (not 100% Neo4j compatible).

• Some advanced graph analytics require custom workarounds.


⸻


🔑 Summary

• FalkorDB = high-performance, Redis-based graph + vector database for real-time AI/ML workloads.

• Best for: recommendation systems, fraud detection, semantic search, knowledge graphs in LLM apps.

• Trade-off: blazing-fast but RAM-heavy and still growing ecosystem compared to Neo4j.


⸻


Tuesday, September 9, 2025

What is Grafitti?

Current RAG approaches struggle when data is updated frequently, limiting their effectiveness for agent-based systems. To solve this, Zep AI’s Graphiti framework introduces a flexible, real-time memory layer built on temporally aware knowledge graphs and stored in Neo4j.

Microsoft’s approach to GraphRAG builds entity-centric knowledge graphs by extracting entities and relationships and grouping them into thematic clusters or “communities.” It relies on LLMs to precompute summaries of these communities. When handling queries, the Microsoft GraphRAG makes multiple LLM calls — first generating partial community-level responses, then combining them into a single comprehensive answer.

Their approach excels at providing detailed, context-rich responses from large static datasets. However, it’s less effective in scenarios where data frequently changes since updates can trigger extensive recomputation of the entire graph. Additionally, its multi-step summarization makes retrieval slow, often taking tens of seconds. This latency and lack of dynamic updates make this approach to GraphRAG unsuitable as a comprehensive and holistic memory for agentic applications

Graphiti With Neo4j: Real-Time Dynamic Agent Memory

Graphiti helps overcome static RAG’s limitations with dynamic data. It’s a real-time, temporally-aware knowledge graph engine that incrementally processes incoming data, instantly updating entities, relationships, and communities without batch recomputation. Graphiti isn’t just another retrieval tool — it’s an ever-present source of context for agents, continuously available and updated.

Graphiti simultaneously handles chat histories, structured JSON data, and unstructured text. All data sources can feed into a single graph, or multiple graphs can coexist within the same Graphiti setup. This gives agents a unified, evolving view of the agent’s world — something traditional RAG systems fundamentally can’t provide.

Why Graphiti Works Better With Dynamic Data

A key feature is Graphiti’s bi-temporal model, which tracks when an event occurred and when it was ingested. Every graph edge (or relationship) includes explicit validity intervals (t_valid, t_invalid). Graphiti uses semantic, keyword, and graph search to determine whether new knowledge conflicts with existing knowledge. When conflicts arise, Graphiti intelligently uses the temporal metadata to update or invalidate, but not discard, outdated information, preserving historical accuracy without large-scale recomputation.

This temporal model enables powerful historical queries, allowing users to reconstruct states of knowledge at precise moments or analyze how data evolves over time.

Fast Query Speeds: Instant Retrieval Without LLM Calls

Graphiti is built for speed. Zep’s own Graphiti implementation achieves extremely low-latency retrieval, returning results at a P95 latency of 300ms. This is enabled by a hybrid search approach that combines semantic embeddings, keyword (BM25) search, and direct graph traversal — avoiding any LLM calls during retrieval.

The use of vector and BM25 indexes offers near-constant time access to nodes and edges, regardless of graph size. This is made possible by Neo4j’s extensive support for both of these index types.

Graphiti’s query latency makes it ideal for real-time interactions, including voice applications.

Graphiti automatically builds an ontology based on incoming data, taking care to de-duplicate nodes and label edge relationships consistently. Beyond automatic ontology creation, Graphiti provides an intuitive method to define custom, domain-specific entities using familiar Pydantic models.

These custom entity types allow precise context extraction, greatly improving the quality of agent interactions. Example entity types might include:

Personalized user preferences and interests (like favorite restaurants, contacts, hobbies), along with standard attributes (name, birthdate, address)

Procedural memory, capturing instructions for actions

Domain-specific business objects (e.g., products, sales orders)

Graphiti automatically matches extracted entities to defined custom types. Custom entity types enhance an agent’s ability to recall knowledge accurately and improve contextual awareness, which is essential for consistent, relevant interactions.


Sunday, September 7, 2025

Kubernetes additional tips - part 2

🔹 3. Copying a Pod for Post-Mortem Debugging


Sometimes a Pod crashes immediately (e.g., due to bad config, missing secrets, startup errors). In such cases, you don’t have enough time to attach a shell or inject an ephemeral container before it dies.


👉 Solution: use kubectl debug --copy-to to clone the Pod into a stable version that won’t crash, so you can perform a post-mortem analysis.


⸻


🔹 Step 1: Clone the Pod into a Debug Version



kubectl debug pod/my-crashing-pod \

  --copy-to=postmortem-pod \

  --image=busybox \

  -- bash -c "sleep 1d"



--copy-to=postmortem-pod → Creates a new Pod called postmortem-pod with the same configuration as the original.

• --image=busybox → Ensures the new Pod uses a stable image with debugging tools (instead of the broken one).

• sleep 1d → Keeps the container running for 1 day (adjust as needed), preventing immediate crash.



Step 2: Exec into the Stable Copy


kubectl exec -it postmortem-pod -- sh


Now you have an interactive shell inside the cloned Pod.


⸻


🔹 What Can You Inspect?


Once inside, you can check:

• Configuration files


cat /etc/config/app.conf


• Mounted secrets


ls /var/run/secrets/kubernetes.io/serviceaccount


Persistent volumes (PVCs) → same mounts as original Pod.

• Environment variables


env | grep DB_


• Application logs left behind in mounted volumes.


This helps you verify if the issue was caused by:

• Wrong config or env vars

• Missing secrets/config maps

• Corrupted volume mounts

• Crash-looping due to command misconfiguration



Why This Is Powerful


✅ Gives you time to debug a Pod that would otherwise crash instantly.

✅ Preserves volumes, secrets, and environment variables for accurate debugging.

✅ Lets you swap the container image for a debug-friendly one (busybox, ubuntu, netshoot, etc.).

✅ Non-destructive — original Pod stays intact (though it may still be crash-looping).



Real-World Example: Debugging a Crashing App


Suppose my-crashing-pod is failing because of a missing DB connection string.

• You clone it with --copy-to.

• Exec in, run env | grep DB_, and discover the variable is not set.

• You check the ConfigMap/Secret mount, realize it’s missing.

• Root cause: the Deployment forgot to mount the db-secret.




 Visual Sequence (Mermaid)


sequenceDiagram

    participant User

    participant kubectl

    participant API[Kubernetes API Server]

    participant PodCrasher[my-crashing-pod (CrashLoopBackoff)]

    participant PodClone[postmortem-pod (Stable copy)]


    User->>kubectl: kubectl debug pod/my-crashing-pod --copy-to=postmortem-pod

    kubectl->>API: Request clone with new image (busybox)

    API->>PodClone: Create postmortem-pod with same config/volumes/env

    User->>PodClone: kubectl exec -it postmortem-pod -- sh

    PodClone->>User: Stable shell for inspection


Summary:

When Pods crash too quickly to debug, cloning them with kubectl debug --copy-to gives you a stable replica for investigation. This allows full inspection of config, volumes, secrets, and logs without modifying the original Pod.



 4. Reading Logs Across Container Restarts


When Pods crash or restart, simply running kubectl logs shows logs from the current running container. This means you miss the previous attempt (which may contain the real cause of the crash).


👉 Kubernetes stores logs for both the current and the last terminated instance of a container.



Viewing the Previous Container’s Logs


# Logs from the last run before restart

kubectl logs my-app-pod -c app --previous


• -c app → If your Pod has multiple containers, specify which one.

• --previous → Fetch logs from the last terminated instance (before the container restarted).


This is essential for debugging CrashLoopBackOff situations, where the Pod dies quickly and restarts.


⸻



 Handling Multiple Restarts


If a Pod restarts many times, you’ll often need logs from each failed run. You can loop through them:


for i in {1..5}; do

  echo "--- Restart #$i ---"

  kubectl logs my-app-pod -c app --previous --since=1h

done


--since=1h → Restricts logs to the last hour (avoids giant logs).

• The loop will repeatedly fetch the previous logs after each restart.


🔑 Note: Kubernetes only retains logs for the last terminated container instance, not the full restart history. For deeper history, you need a log aggregator (e.g., EFK/ELK, Loki, Datadog, Splunk).


 Trimming Large Logs


When containers generate huge logs, --since and --tail are lifesavers:


# Last 100 lines from the previous instance

kubectl logs my-app-pod -c app --previous --tail=100


# Logs from the last 30 minutes

kubectl logs my-app-pod -c app --previous --since=30m


 Debugging Workflow

. Check Pod status & restarts


kubectl get pod my-app-pod

kubectl describe pod my-app-pod


Look at Restart Count and termination reasons.


2. Fetch logs from the last crash


kubectl logs my-app-pod -c app --previous


3. Filter for timeframe or lines if logs are too large.

4. Escalate to log aggregation if you need full history beyond one restart.



Visual Flow (Mermaid)

sequenceDiagram

    participant User

    participant kubectl

    participant KubeAPI

    participant Pod

    participant Container


    User->>kubectl: kubectl logs my-app-pod -c app

    kubectl->>KubeAPI: Request logs (current container)

    KubeAPI->>Container: Fetch running logs

    Container->>User: Returns logs (current only)


    User->>kubectl: kubectl logs my-app-pod -c app --previous

    kubectl->>KubeAPI: Request logs (terminated container)

    KubeAPI->>Pod: Get last restart logs

    Pod->>User: Returns logs from previous instance


 Summary:

• kubectl logs → current instance logs.

• kubectl logs --previous → last terminated instance logs (great for crash debugging).

• Use --since / --tail to keep logs manageable.

• For full restart history, integrate a centralized logging solution (ELK, Loki, etc.).


⸻



Would you like me to also show how to pipe logs directly into grep/jq for structured filtering (e.g., finding error stack traces across restarts)?


Kubernetes additional tips - Part 1

 1. Understanding kubectl port-forward


By default, when you use kubectl port-forward, it only binds to 127.0.0.1 (localhost).

• That means only processes on your own machine can reach the forwarded port.

• This is safer because it prevents external access unless you explicitly allow it.


⸻


🔹 Example: Forwarding Pod Port to Local Machine


Suppose your Pod (my-app-pod) exposes port 80, and you want to access it locally on 8080:


kubectl port-forward pod/my-app-pod 8080:80


Maps Pod’s port 80 → localhost:8080

• Accessible only on your machine, not from the network.


⸻


🔹 Exposing Beyond Localhost


If you want to allow other machines on your network (LAN or cloud VM) to access it:


kubectl port-forward --address 0.0.0.0 pod/my-app-pod 8080:80


--address 0.0.0.0 → binds to all network interfaces.

• Now anyone who can reach <your-ip>:8080 can access the Pod’s service.


✅ Useful for:

• Testing webhooks from external services

• Allowing teammates to test your Pod’s API or UI

• Debugging cluster apps from another device




 Security Considerations

• Opening port-forward to 0.0.0.0 is not secure on untrusted or public networks.

• Anyone with access to your IP can now hit your Pod directly.

• Mitigation options:

• Restrict with firewall rules (e.g., allow only certain IP ranges).

• Use SSH tunnels instead of --address 0.0.0.0.

• Prefer creating a Kubernetes Service + Ingress for controlled external access.


⸻


🔹 Pro Tip: Port-Forwarding a Service Instead of a Pod


Instead of binding to a specific Pod (which may restart or move), you can forward a Service:



2. Injecting an Ephemeral Debug Container in Kubernetes


Sometimes your Pod’s image is minimal (like nginx, httpd, or Google’s distroless images) and does not include debugging tools (e.g., no bash, sh, curl, ps).

Instead of rebuilding your image, Kubernetes lets you attach an ephemeral debug container into the running Pod, sharing its namespaces (network, process, filesystem).


Ephemeral containers were introduced in Kubernetes v1.18 (beta by v1.23, stable in v1.25).


Step 1: Create the Base Pod (without debug tools)


apiVersion: v1

kind: Pod

metadata:

  name: debug-demo

spec:

  containers:

  - name: app

    image: nginx


kubectl apply -f debug-demo.yaml


At this point, the Pod runs nginx but has no shell or debugging tools.


🔹 Step 2: Inject an Ephemeral Container


kubectl debug debug-demo \

  --image=busybox \

  --target=app \

  --name=debug-shell \

  -- bash


• --image=busybox → use BusyBox (has sh, curl, netcat, etc.). You could also use ubuntu, debian, or nicolaka/netshoot for richer debugging tools.

• --target=app → joins the namespaces (network, process, IPC, etc.) of the main container (app).

• --name=debug-shell → names the ephemeral container (must be unique in the Pod).

• bash → command to run (if the image has only sh, replace bash with sh).


You’ll be dropped into a live shell with full access to the Pod’s filesystem, network, and process space.



 What Can You Do Inside?


Ephemeral containers let you observe and debug without modifying your production image:

• Inspect filesystems


ls /usr/share/nginx/html

cat /etc/nginx/nginx.conf



Test connectivity


curl localhost:80

nc -vz other-service 3306



Check processes


ps aux

strace -p 1


Debug networking with netshoot (traceroute, tcpdump, dig, etc.).


Key Points & Best Practices

• Ephemeral containers are temporary — they don’t survive Pod restarts or rescheduling.

• They don’t modify PodSpec in a way that affects restarts — they’re strictly for debugging.

• They are not restarted automatically if they exit.

• Useful for:

• Debugging distroless or minimal images

• Inspecting live network traffic inside a Pod

• Running diagnostics tools without polluting production images

• Security: RBAC permissions (ephemeralcontainers) must be enabled for your user.


sequenceDiagram

    participant User

    participant kubectl

    participant KubeAPI

    participant Pod[Pod: nginx app container]

    participant DebugContainer[Ephemeral Container: busybox]


    User->>kubectl: kubectl debug debug-demo --image=busybox --target=app

    kubectl->>KubeAPI: Request to inject ephemeral container

    KubeAPI->>Pod: Add temporary debug container

    Pod->>DebugContainer: Start busybox (bash/sh)

    User->>DebugContainer: Gets shell access

    DebugContainer->>Pod: Shares namespaces (network, FS, processes)

    User->>DebugContainer: Run tools (curl, ps, strace)




Summary:

Ephemeral debug containers are the safest, fastest way to troubleshoot live Pods without rebuilding or redeploying images. They let you inject debugging tools on demand, inspect traffic and files, and troubleshoot issues inside the Pod’s environment.


⸻


Would you like me to also show you a real-world example using nicolaka/netshoot (which comes preloaded with dozens of network debugging tools) instead of just busybox?



Parent ref 

https://medium.com/@obaff/7-kubernetes-debugging-tricks-you-probably-didnt-know-832884bd90a1


Kubernetes: high level best practices

1. Running Everything in the Default Namespace

This quickly becomes messy, especially when you scale to multiple apps and environments.

Fix: Always create and use separate namespaces for applications, environments (dev, staging, prod), and monitoring tools.

2. Forgetting to Define Resource Requests and Limits

Without CPU and memory requests/limits, pods can consume unlimited resources, leading to instability or even node crashes.

Fix: Always define resources.requests and resources.limits in your pod specs.

3. Ignoring Liveness and Readiness Probes

Beginners often skip health probes. Without them, Kubernetes won’t know when your app is ready or if it’s stuck.

 Fix: Always configure readiness probes (for traffic) and liveness probes (for restarts).

4. Hardcoding Configurations Inside Pods

Putting configuration values (like DB passwords, API keys) directly inside pod definitions is a rookie mistake.

Fix: Use ConfigMaps for non-sensitive configs and Secrets for sensitive data.

5. Exposing Applications Using NodePort

Many start with NodePort, but it’s not production-grade and makes apps hard to access securely.

Fix: Use Ingress controllers (NGINX, Traefik, etc.) with proper domain names and TLS.

6. Not Using Labels and Selectors Properly

Labels are the glue of Kubernetes. Without consistent labeling, managing workloads, deployments, and monitoring is a nightmare.

Fix: Define a clear labeling strategy (app, env, version) and stick to it.

7. Overlooking RBAC (Role-Based Access Control)

Running everything with cluster-admin privileges is risky. It’s common for beginners to skip RBAC setup entirely.

Fix: Use least-privilege access and set up RBAC roles early.

8. Forgetting About Persistent Volumes

Beginners often assume storage works like stateless pods. When a pod restarts, all data inside disappears.

 Fix: Use PersistentVolumes (PV) and PersistentVolumeClaims (PVC) for stateful apps.

9. Not Monitoring and Logging

Kubernetes without observability is like flying blind. Many beginners only check pod status and logs manually.

Fix: Use monitoring tools like Prometheus + Grafana and logging with ELK stack or Loki.

10. Deploying Without Understanding the Basics

Many jump straight into complex Helm charts and operators without understanding Pods, Services, and Deployments first.

Fix: Master the fundamentals before moving on to advanced tooling.

references:
https://aws.plainenglish.io/10-kubernetes-mistakes-beginners-make-and-how-to-avoid-them-5c92f5766605

Monday, September 1, 2025

What is Zep GraphRAG memory?

 Zep’s Graph RAG is a dynamic, temporally-aware retrieval system built on top of a continuously updated knowledge graph (Graphiti). Unlike standard RAG (Retrieval-Augmented Generation) that usually works with static documents, Zep’s Graph RAG is designed to handle real-time business and conversational data—including customer interactions, support tickets, preferences, and more—without expensive batch recomputation.


Additional Claims & Strengths of Zep’s Approach

• Continuous Data Integration: Ingests JSON, text, chat, or structured business data in real time—it immediately becomes part of the knowledge graph.



• Temporal Fact Management: Keeps a historical timeline of facts; knows what’s current vs. invalid—helping agents reason about evolving situations.

  

• Relationship-Aware Retrieval: Retrieves context across related entities—e.g., given a customer, fetch their support tickets, purchases, preferences, etc.



• Shared Knowledge Graphs: Supports graphs shared across users or agents for domain-wide context—centralized knowledge storage.



• Custom Entity Types: Developers can model domain-specific entities (customers, products, projects) and define relationships relevant to their business.

 

• Low-Latency, Scalable: Handles large-scale datasets at low latencies (<ms), combining multiple retrieval strategies.

 

• Temporal Agents Memory Layer: Built for agent memory—not just RAG. The architecture, powered by Graphiti, handles both unstructured conversation and structured data with temporal insights.

 


⸻


In Summary:


Zep’s Graph RAG Memory evolves standard RAG by making it:

• Dynamic (real-time updates)

• Temporal (historical context and fact validity)

• Fast (millisecond-level retrieval)

• Context-rich (relationship-aware and domain-aware)


This enables AI agents to retrieve up-to-date, nuanced, temporal context rather than static snapshots—improving accuracy, reducing hallucinations, and allowing smarter decisioning.


Would you like me to show you a specific API example or a developer demo of how this is used in code?

What is Graphiti

Graphiti helps overcome static RAG’s limitations with dynamic data. It’s a real-time, temporally-aware knowledge graph engine that incrementally processes incoming data, instantly updating entities, relationships, and communities without batch recomputation. Graphiti isn’t just another retrieval tool — it’s an ever-present source of context for agents, continuously available and updated.

Graphiti’s real-time incremental architecture is built for frequent updates. It continuously ingests new data episodes (events or messages), extracting and immediately resolving entities and relationships against existing nodes.

A key feature is Graphiti’s bi-temporal model, which tracks when an event occurred and when it was ingested. Every graph edge (or relationship) includes explicit validity intervals (t_valid, t_invalid). Graphiti uses semantic, keyword, and graph search to determine whether new knowledge conflicts with existing knowledge. When conflicts arise, Graphiti intelligently uses the temporal metadata to update or invalidate, but not discard, outdated information, preserving historical accuracy without large-scale recomputation.

Fast Query Speeds: Instant Retrieval Without LLM Calls

Graphiti is built for speed. Zep’s own Graphiti implementation achieves extremely low-latency retrieval, returning results at a P95 latency of 300ms. This is enabled by a hybrid search approach that combines semantic embeddings, keyword (BM25) search, and direct graph traversal — avoiding any LLM calls during retrieval.

Graphiti represents a meaningful departure from traditional RAG methods, specifically because it was built from the ground up as a memory infrastructure for dynamic agentic systems. Graphiti offers incremental, real-time updates through its temporally aware knowledge graph. This design means engineers no longer need to recompute entire graphs when data changes. Instead, Graphiti incrementally integrates updates, resolves conflicts based on temporal metadata, and maintains an accurate historical state.

By removing the bottleneck of LLM-driven summarization at query time, Graphiti achieves practical latency levels that engineers require for interactive real-world applications. Its hybrid indexing system — combining semantic embeddings, keyword search, and graph traversal — allows rapid retrieval in near-constant time, independent of graph scale. With intuitive tools like custom entity types implemented through familiar structures such as Pydantic models, Graphiti addresses a significant capability gap in agent development, equipping engineers with a robust, performant, and genuinely dynamic memory layer.