← posts / ai integrations

Steal Palantir's agent stack: typed tools, one LLM gateway, swappable models

An X thread boils Palantir's AIP docs down to four agent patterns. We check each one against the docs, then build them in one stdlib-only Python file: typed business-object tools, a gateway that masks PII, caches and retries, a model set in config, and schedule/event/API triggers.

if.codesOct 1, 2026 · 10 min read#ai-agents#llm#python#architecture#palantirAI-assisted

A thread by @undefinedKi reads Palantir's AI platform docs and pulls out four patterns you can copy into your own project: the agent sees business objects instead of raw data, every LLM call goes through one gateway, the model name lives in config, and agents start from a schedule, an event, or an API call.

That's a good list. This post checks each point against Palantir's public docs, so you can see which parts are Palantir's and which are the thread's advice. Then it builds all four in one stdlib-only Python file (about 230 lines) that runs without an API key.

What Palantir's docs actually say

1. Typed business objects and actions, not tables

The Ontology overview says data sources get mapped into "objects, properties, and links," and that change happens through "action types and functions." The AIP agent tools page lists the tools an agent can be given. These include Object Query (filter, aggregate and traverse links on the object types you pick), Action (make an Ontology edit, optionally only after the user confirms) and Function. You pick which object types and actions each agent can reach, and the model never writes SQL.

A conceptual representation of a business ontology replacing raw data tables.
A conceptual representation of a business ontology replacing raw data tables.

2. One governed path to the model

This is where you need to be careful about what Palantir actually says. Its docs describe LLM-provider-compatible proxy endpoints that "benefit from Foundry capabilities such as rate limiting, data governance, and usage tracking," and that enforce zero data retention and georestriction. LLM capacity management sets token-per-minute and request-per-minute limits per enrollment, per project and per user. The AIP security page says third-party providers don't retain prompts or completions. For caching, Pipeline Builder's Use LLM node can skip rows whose prompt inputs haven't changed since an earlier build.

A metaphor for a single, governed gateway controlling all access to a model.
A metaphor for a single, governed gateway controlling all access to a model.

I couldn't find a Palantir doc that says prompts get PII-masked on their way out, and none that describes an automatic retry policy. Both are the thread's suggestions, and they're good ones. Palantir does sell a Sensitive Data Scanner, but that's a separate product.

3. Swappable models

Supported LLMs covers models from xAI, OpenAI, Anthropic, Meta and Google. Bring your own model lets you register an externally hosted or self-hosted LLM for use in AIP Logic, Chatbot Studio and other apps. In the thread's words, keep the model name in config, and switching becomes a one-line change.

An illustration of swappable model components in a modular system.
An illustration of swappable model components in a modular system.

4. Schedule, event, API

Automate supports time conditions (hourly, daily, weekly, monthly, or custom cron) and object set conditions (objects added, removed or modified, threshold crossed). An automation can call a Logic function for each new object. For the API side, the Logic getting-started guide shows copying a curl command to run a function from outside Foundry (not for Logics that return Ontology edits), and published queries can be called through the Execute Query API.

The minimal version

Save this as agent_stack.py. It needs Python 3.10+ and nothing else: no pip install, no key. SQLite stands in for your database, a fake provider stands in for the LLM, and MODEL picks the provider.

"""agent_stack.py - a minimal "ontology + gateway + triggers" agent stack.

Stdlib only (Python 3.10+). Run:
    MODEL=fake python3 agent_stack.py          # event + schedule demo
    MODEL=fake python3 agent_stack.py serve    # API trigger on :8765
"""
import hashlib, json, os, re, sched, sqlite3, sys, time, urllib.error, urllib.request
from dataclasses import dataclass, asdict
from http.server import BaseHTTPRequestHandler, HTTPServer

# ---------------------------------------------------------------- 1. data
DB = sqlite3.connect(":memory:", check_same_thread=False)
DB.executescript("""
CREATE TABLE customers(id TEXT PRIMARY KEY, name TEXT, email TEXT, phone TEXT, tier TEXT);
CREATE TABLE orders(id TEXT PRIMARY KEY, customer_id TEXT, total REAL, status TEXT);
INSERT INTO customers VALUES ('c1','Ada Lovelace','ada@example.com','+44 20 7946 0958','gold');
INSERT INTO orders VALUES ('o1','c1',120.0,'new'), ('o2','c1',4800.0,'new');
""")

# ---------------------------------------------------------------- 2. business objects
@dataclass(frozen=True)
class Customer:
    id: str
    name: str
    email: str
    phone: str
    tier: str

@dataclass(frozen=True)
class Order:
    id: str
    customer_id: str
    total: float
    status: str

# ---------------------------------------------------------------- 3. typed tools
TOOLS = {}

def tool(fn):
    """Register a function as an agent tool. The LLM only ever sees these."""
    TOOLS[fn.__name__] = fn
    return fn

@tool
def get_order(order_id: str) -> Order:
    row = DB.execute("SELECT id, customer_id, total, status FROM orders WHERE id=?",
                     (order_id,)).fetchone()
    if row is None:
        raise LookupError(f"order {order_id} not found")
    return Order(*row)

@tool
def get_customer(customer_id: str) -> Customer:
    row = DB.execute("SELECT id, name, email, phone, tier FROM customers WHERE id=?",
                     (customer_id,)).fetchone()
    if row is None:
        raise LookupError(f"customer {customer_id} not found")
    return Customer(*row)

@tool
def flag_order_for_review(order_id: str, reason: str) -> dict:
    """An *action*: the only write the agent is allowed to make."""
    if not reason.strip():
        raise ValueError("reason is required")
    DB.execute("UPDATE orders SET status='review' WHERE id=? AND status='new'", (order_id,))
    return {"order_id": order_id, "status": "review", "reason": reason}

def call_tool(name: str, args: dict) -> dict:
    if name not in TOOLS:
        raise PermissionError(f"tool {name!r} is not on the allow-list")
    result = TOOLS[name](**args)
    return asdict(result) if hasattr(result, "__dataclass_fields__") else result

def tool_catalog() -> str:
    return "\n".join(f"- {n}{TOOLS[n].__annotations__}" for n in TOOLS)

# ---------------------------------------------------------------- 4. providers
class TransientError(Exception):
    """Timeouts, 429s, 5xx: worth retrying."""

class FakeProvider:
    """Deterministic stand-in so the demo runs with no API key.
    Hard-coded to this file's prompt format (it looks for TOOL/order
    lines); a real provider replaces it, not extends it.
    Fails its first call to show the retry path."""
    def __init__(self):
        self.calls = 0
        self.last_prompt = ""

    def complete(self, model: str, prompt: str) -> str:
        self.calls += 1
        self.last_prompt = prompt
        if self.calls == 1:
            raise TransientError("simulated 503")
        seen = re.findall(r"TOOL (\w+) -> (.*)", prompt)
        done = {name: json.loads(body) for name, body in seen}
        task = re.search(r"order (\w+)", prompt).group(1)
        if "get_order" not in done:
            return json.dumps({"tool": "get_order", "args": {"order_id": task}})
        order = done["get_order"]
        if "get_customer" not in done:
            return json.dumps({"tool": "get_customer",
                               "args": {"customer_id": order["customer_id"]}})
        if order["total"] > 1000 and "flag_order_for_review" not in done:
            return json.dumps({"tool": "flag_order_for_review",
                               "args": {"order_id": order["id"],
                                        "reason": "total above 1000"}})
        return json.dumps({"answer": f"order {order['id']} checked, total {order['total']}"})

class OpenAICompatibleProvider:
    """Any /v1/chat/completions endpoint (OpenAI, OVH, vLLM, Ollama, LiteLLM proxy...)."""
    def __init__(self):
        self.base = os.environ.get("LLM_BASE_URL", "https://api.openai.com/v1")
        self.key = os.environ["LLM_API_KEY"]

    def complete(self, model: str, prompt: str) -> str:
        req = urllib.request.Request(
            f"{self.base}/chat/completions",
            data=json.dumps({"model": model,
                             "messages": [{"role": "user", "content": prompt}]}).encode(),
            headers={"Authorization": f"Bearer {self.key}",
                     "Content-Type": "application/json"})
        try:
            with urllib.request.urlopen(req, timeout=60) as r:
                return json.load(r)["choices"][0]["message"]["content"]
        except urllib.error.HTTPError as e:
            if e.code == 429 or e.code >= 500:
                raise TransientError(str(e)) from e
            raise
        except (urllib.error.URLError, TimeoutError) as e:    # timeouts, refused connections
            raise TransientError(str(e)) from e

# ---------------------------------------------------------------- 5. the gateway
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
PHONE = re.compile(r"\+?\d[\d\s().-]{7,}\d")

def mask_pii(text: str) -> str:
    return PHONE.sub("[PHONE]", EMAIL.sub("[EMAIL]", text))

class Gateway:
    """Every LLM call in the app goes through here - nowhere else."""
    def __init__(self):
        spec = os.environ.get("MODEL", "fake")           # e.g. "fake" or "openai:gpt-4o-mini"
        provider, _, self.model = spec.partition(":")
        self.provider = {"fake": FakeProvider,
                         "openai": OpenAICompatibleProvider}[provider]()
        self.max_retries = int(os.environ.get("LLM_MAX_RETRIES", "3"))
        self.cache: dict[str, str] = {}
        self.stats = {"calls": 0, "cache_hits": 0, "retries": 0}

    def complete(self, prompt: str) -> str:
        prompt = mask_pii(prompt)                          # mask before it leaves
        key = hashlib.sha256(f"{self.model}\x00{prompt}".encode()).hexdigest()
        if key in self.cache:
            self.stats["cache_hits"] += 1
            return self.cache[key]
        for attempt in range(self.max_retries + 1):
            try:
                self.stats["calls"] += 1
                out = self.provider.complete(self.model, prompt)
                self.cache[key] = out
                return out
            except TransientError:
                if attempt == self.max_retries:
                    raise
                self.stats["retries"] += 1
                time.sleep(0.1 * 2 ** attempt)             # 0.1s, 0.2s, 0.4s ...
        raise RuntimeError("unreachable")

GATEWAY = Gateway()

# ---------------------------------------------------------------- 6. the agent loop
def run_agent(task: str, max_steps: int = 6) -> str:
    transcript = [f"TASK {task}", "Reply with JSON: {\"tool\",\"args\"} or {\"answer\"}.",
                  "TOOLS:\n" + tool_catalog()]
    for _ in range(max_steps):
        reply = json.loads(GATEWAY.complete("\n".join(transcript)))
        if "answer" in reply:
            return reply["answer"]
        result = call_tool(reply["tool"], reply.get("args", {}))
        transcript.append(f"TOOL {reply['tool']} -> {json.dumps(result)}")
    raise RuntimeError("agent did not finish")

# ---------------------------------------------------------------- 7. triggers
SUBSCRIBERS: dict[str, list] = {}

def on_event(name):
    def register(fn):
        SUBSCRIBERS.setdefault(name, []).append(fn)
        return fn
    return register

def publish(name, payload):
    for fn in SUBSCRIBERS.get(name, []):
        fn(payload)

@on_event("order.created")
def review_new_order(payload):
    print("[event]   ", run_agent(f"review order {payload['order_id']}"))

def nightly_sweep():
    for (oid,) in DB.execute("SELECT id FROM orders WHERE status='new'").fetchall():
        print("[schedule]", run_agent(f"review order {oid}"))

class API(BaseHTTPRequestHandler):
    def do_POST(self):
        try:
            length = int(self.headers.get("Content-Length") or 0)
            body = json.loads(self.rfile.read(length) or b"{}")
            status, out = 200, {"answer": run_agent(f"review order {body['order_id']}")}
        except (ValueError, KeyError, TypeError, LookupError) as e:
            status, out = 400, {"error": str(e)}
        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.end_headers()
        self.wfile.write(json.dumps(out).encode())

if __name__ == "__main__":
    if len(sys.argv) > 1 and sys.argv[1] == "serve":
        print("POST http://127.0.0.1:8765  {\"order_id\": \"o1\"}")
        HTTPServer(("127.0.0.1", 8765), API).serve_forever()
        sys.exit(0)
    publish("order.created", {"order_id": "o2"})          # event trigger
    s = sched.scheduler(time.monotonic, time.sleep)      # schedule trigger
    s.enter(0.2, 1, nightly_sweep)                       # in prod: cron / APScheduler
    s.run()                                              # blocks until the queued job has run
    publish("order.created", {"order_id": "o2"})          # same task again -> cache
    print("statuses:", DB.execute("SELECT id, status FROM orders").fetchall())
    print("gateway: ", GATEWAY.stats)
    print("provider saw:", [l for l in GATEWAY.provider.last_prompt.splitlines()
                            if "get_customer" in l and "TOOL" in l])

Run it

MODEL=fake python3 agent_stack.py

MODEL defaults to fake, so plain python3 agent_stack.py works too (on Windows, where the VAR=value command prefix isn't supported, use that or set MODEL=fake first). Output on my machine (Python 3.14):

[event]    order o2 checked, total 4800.0
[schedule] order o1 checked, total 120.0
[event]    order o2 checked, total 4800.0
statuses: [('o1', 'new'), ('o2', 'review')]
gateway:  {'calls': 11, 'cache_hits': 1, 'retries': 1}
provider saw: ['TOOL get_customer -> {"id": "c1", "name": "Ada Lovelace", "email": "[EMAIL]", "phone": "[PHONE]", "tier": "gold"}']

Each line shows one of the patterns working:

  • Typed tools: the agent read an Order and a Customer and moved o2 into review through the one write action it's allowed. It never saw a table name or a SQL string.
  • PII masking: the provider saw line is the exact tool result the model received. The email and phone number were replaced before the prompt left the process.
  • Retry: the fake provider fails its first call with a simulated 503. The gateway backed off, tried again, and counted it ('retries': 1).
  • Cache: the second order.created event started with the same prompt as the first, so that step came from the cache ('cache_hits': 1). Later steps missed because o2's status had changed, and that's what you want. The cache key includes the model name.
  • Triggers: the same run_agent function ran from an event and from a scheduled sweep. The sweep skipped o2 because it was no longer new.

For the API trigger, start the server and POST to it:

MODEL=fake python3 agent_stack.py serve
# in another terminal
curl -s -X POST localhost:8765 -d '{"order_id":"o1"}'
# {"answer": "order o1 checked, total 120.0"}
Want AI wired into the systems you already run?I build LLM integrations with costs and quality you can see. The estimate is free.

Plug in a real model

The OpenAICompatibleProvider class talks to any /chat/completions endpoint: OpenAI, a LiteLLM proxy, vLLM, or Ollama's OpenAI-compatible API. Switching is an environment change:

export MODEL=openai:gpt-4o-mini
export LLM_API_KEY=...          # your key
export LLM_BASE_URL=https://api.openai.com/v1   # or http://localhost:11434/v1 for Ollama
python3 agent_stack.py

This path wasn't run for this post (no key was used), and a real model may not follow the JSON reply format as closely as the fake one does. Before you ship it, add a parse-and-retry step, or use your provider's native tool-calling or JSON mode. To add another provider, write a class with complete(model, prompt) -> str, raise TransientError on 429s and 5xx, and add it to the dict in Gateway.__init__.

What to harden before production

  • Masking: two regexes miss names, addresses and IDs. Better options are to drop fields in the tool layer (the model rarely needs email) or to use a dedicated detector such as Microsoft Presidio.
  • Cache: an in-process dict disappears on restart and isn't shared between workers. Move it to Redis with a TTL, and only cache calls whose answers can safely repeat.
  • Actions: the Palantir pattern of asking the user to confirm before an edit is easy to copy. Have write tools return a pending change that a person approves.
  • Accounting: the gateway already sees every call, so it's where token counts, per-user limits and an audit log should go. That matches what the Palantir docs describe for their proxy endpoints.
  • Triggers: swap sched for cron or APScheduler, the in-memory bus for Kafka, NATS or Redis Streams, and http.server for your web framework. run_agent doesn't need to change.

The shape itself is what matters: a short list of typed tools, one function that owns every model call, and a config value that decides which model answers.

Sources

if.codesI build RAG, AI integrations and agent pipelines on Go and Python backends — and write about it here.
// keep reading

More posts on AI and backends.

// free quote

Read something you need? I’ll quote it for free.

RAG, AI integrations, agents or the backend underneath — tell me what you have and what should change. I read every request myself.

Free · no commitment

Tell me what you have. I’ll tell you what it takes.

1Describe the projectA few sentences is enough — about two minutes.
2I review itI read it myself and may ask a follow-up question.
3You get a free quoteScope, approach and estimate — yours to keep, no strings.
What kind of project is it?
Free and without obligation. Your details are used only to reply — see the privacy policy.