Open Trading Surface

User Guide

How Open Trading Surface works, how to set it up, and how to use every part of it. If you're new, start with the 5-minute quickstart, then dip into whichever feature you need. If you last used it before the panel workspace arrived, read The terminal is a workspace first — the way you move around has changed.

If you are coming back after a while, four things arrived that change how you work rather than merely adding to it: the economic calendar and the one series identity underneath it (including a dated change to how filed fundamentals are dated); Perspectives becoming three lenses with a prompt library of its own; the thesis engine gaining an AI discovery agent and a four-stage journey; and chart preferences — the defaults every chart opened with are now yours to set.
Getting started
★ 5-minute quickstart How Open Trading Surface works Requirements Installing Configuring API keys First run & the disclaimer Automations & scheduling
Using the terminal
The terminal is a workspace Driving it with the agent Administration Processes One selector, two-tier disclosure
Research
Charts & the workbench Screening stocks Modernized G & D Cohorts Pricing engine & Model of Models Company operating models Pro forma statements & the substrate The economy & industry containers The economic calendar Series, overlays & knowability Perspectives — the three lenses The thesis engine & discovery The Market tab: heat map & repricing scan Divergence ledger, reratings & event impact Honest numbers: grades, base rates & expectations
Community — shared research
The Community tab: chats & channels Sharing research, and the hive mind Why sharing is safe here
Portfolios
Portfolios & the transaction ledger Sorting & the house table Cost basis & tax lots Performance, attribution & risk The PDF report Where a holding came from
Event processing
→ The full event-processing guide Reach: what can honestly be backtested Typed events The condition language Rules & the tick Actions & the paper book Strategies The backtester & deflation Events, Firings, Runs & Backtests
Everything else
Alerts, subscriptions & the calendar Executive profiles The marketplace & images Research notes & memory Giving feedback Updating Uninstalling Your data & privacy Troubleshooting

★ 5-minute quickstart

  1. Install the plugin and restart your Claude session (details below).
  2. Add your FMP key: say "open the Open Trading Surface terminal", click 🛠 Administration in the top strip, pick Settings from the left menu, paste your key, press Test.
  3. Accept the one-time disclaimer when prompted.
  4. Try your first real task — just ask the agent in plain language:
    "Analyze NVDA — chart it with the 50/200-day averages and RSI,
     then give me the fundamental picture and what the bulls and bears would say."
  5. Then explore one thing: "screen large-cap tech under 20× earnings ranked by momentum", or "open a chart on AAPL and a chart on MSFT side by side."
  6. Turn on the batches when you're ready for the twin to accrue: run /ots:automate and pick the scheduled jobs you want (see Automations).
You don't click Open Trading Surface like a dashboard — you tell the agent what you want and it drives the terminal, panels, charts, screens, and models for you. The sections below show what's possible and the phrasings that work.

How Open Trading Surface works

Two ideas explain everything, and a third explains its character.

1. It's a price-discovery digital twin, not a data feed

The market is a price discovery mechanism, not a forecast — so Open Trading Surface doesn't try to out-predict it. It maintains a digital twin that discovers an evidence-supported price per name, reconciles it against the market's price, and treats the unexplained gap — the residual — as the object of study: mapped market-wide on the heat map, logged in the divergence ledger, and tested for reratings. Around that core runs the connected chain: a worldview (macro, geopolitics, demographics) becomes theses with explicit, auditable probability estimates and an evidence ledger; theses compile into screens, cohorts and paper portfolios; individual names are grounded in driver-based operating models, an articulated pro forma statement set, and a mixture of valuation frames. It persists between sessions and is continuously reassessed as evidence arrives. Everything you build compounds instead of resetting.

2. It's agent-native

The terminal is a shared world-model: everything you see, the agent can read and drive. You work by conversation — "open a chart on AAPL beside MSFT," "what do the signals say," "record this as a thesis." The agent inspects the exact panel you're looking at and acts on it. It's plan-only: it models, screens, and drafts, but never places a trade or moves money.

3. It refuses rather than defaults

This is the property that distinguishes it, and it is deliberate. Where a number cannot be honestly computed, Open Trading Surface says so by name instead of substituting a plausible one — which input is missing, on which leg, and what would fix it. Where a new layer would move a number you have already read, it discloses before it adjusts: both values side by side, the delta named, and the change recorded as a dated methodology break. And it will tell you what it could honestly backtest rather than producing a flattering number over a window it does not have the record for. Empty results, red process lights and refusal sentences are not failures of the product — they are the product working.

The clearest recent example: filed fundamentals used to be treated as knowable on the day the period ended rather than on the day the filing that published them existed. Fixing that moved published numbers by a median of 496 days — so the release that disclosed the dating basis shipped first, moving no number at all, and the correction followed. A disclosure no reader carries is a comment.

About the screenshots on this page. Every image here was captured from a live terminal on 2026-08-07. Where the surface has moved since, the caption says what moved rather than letting the picture stand as current. Several surfaces described below — the economic calendar, the thesis journey, the prompt library, chart preferences — have no screenshot at all yet, and none has been invented for them.
Open Trading Surface is not investment advice. Every output — including probabilities, models, fair values, characterizations and paper-book proposals — is a decision aid to verify independently. You bear full responsibility for all decisions.

Requirements

Installing

Claude desktop (Cowork)

Open Trading Surface is a plugin. Install it from the Customize menu:

  1. Download the latest package — ots.plugin.
  2. Open the Cowork tab, then open Customize in the left sidebar and go to the Plugins tab. (Customize is where plugins, skills, and connectors now live.)
  3. Under Personal plugins, click the + button and add the ots.plugin file you downloaded. Plugins you add yourself are stored locally on your computer.
  4. Restart the session so the plugin's local server starts. Open Trading Surface's skills then appear under the / menu, and its tools as mcp__ots__*.
  5. You should see no format warning. Through 5.52.0 the package shipped its entry points in the older commands/ layout and installing it printed a note suggesting the skills/ format. From 5.53.0 every entry point is a skill, so that note is gone. If you still see it, you are installing an older package.
  6. Start it: in a Cowork chat, tell the agent "open the Open Trading Surface terminal." That boots Open Trading Surface's local engine and opens the terminal in your browser — the first thing to do after installing.
Nothing to launch manually — you never run a command yourself. Installing the plugin makes the tools available; saying "open the Open Trading Surface terminal" (or just "open Open Trading Surface") is what actually starts it.
The / skills are yours to trigger, not Claude's. Every one of them — /ots:analyze, /ots:chart, /ots:compare, /ots:fundamentals, /ots:screen, /ots:setup, /ots:automate — is declared explicit-invocation only, so Claude will never decide to run one on its own. That is deliberate: two of them write. /ots:setup stores an API key, and /ots:automate installs scheduled tasks on your machine, and neither should happen because a sentence read as though it were a request. The single exception is the stock-analysis methodology skill, which Claude may consult when you ask about a ticker — it describes how to reason and touches nothing.

Configuring API keys

The easiest path is in the UI: open the terminal, click 🛠 Administration in the top strip, and pick Settings from the left menu. That one panel holds every credential the system uses:

You can also set FMP_API_KEY / FRED_API_KEY / OTS_AI_API_KEY environment variables, which take precedence over the stored file.

Keys are stored on your machine in ~/.ots/config.json with file mode 600 and are never transmitted anywhere but the data providers themselves. Exactly one AI vendor is configured at a time — saving a different one replaces the configuration rather than parking the old key beside it. There is no simultaneous-vendor mode and no automatic failover, deliberately.
What a Test button establishes, and what it does not. The EDGAR check reports that the request was accepted, never that your contact string was approved — measured on 2026-08-10, a deliberately fabricated contact returned 200, so SEC does not enforce identity at the wire on that endpoint today and a tick mark implying otherwise would manufacture an authority nobody granted. The local Name email@domain.tld shape check is ours, says so, and never overrides the network verdict.

Prompts are not in Settings. Every text this installation sends to a model — the three lens defaults, per-security overrides and library, and the thesis-discovery prompt — lives in Administration → Prompts; see Perspectives.

First run & the disclaimer

On first launch the terminal presents a one-time legal acknowledgment. You must read and accept it before using the application. A short disclaimer stays visible in the footer at all times, and the full terms live on the Legal page.

Open Trading Surface is not investment advice. Every output is a decision aid to be independently verified. You bear full responsibility for all investment decisions.

Automations & scheduling

The twin accrues by running batches. The plugin ships the capabilities, but scheduled tasks live on your machine, so installing the plugin does not install them. Run /ots:automate and the agent walks the registry with you: for each automation it shows the title, what it does, the suggested cadence, and a heads-up for the heavy ones, then creates only the tasks you choose.

Whatever you install, Administration → Processes is where you see whether it actually ran. If your environment has no scheduling capability, /ots:automate gives you the command-line equivalent for the automations that have one, and says plainly which have no headless equivalent.

Automations are read-only over your configuration. They never modify strategic-initiative books, pricing models, watchlists or settings; everything they record is descriptive market data.

The terminal is a workspace

The right-hand area of the terminal is no longer one active view — it is an MDI floating panel workspace. You can open any number of panels, arrange them freely, and reduce them without closing them.

The terminal workspace with six panels tiled three across and two down on a single security, each panel carrying its own title bar with duplicate, minimize, maximize and close controls, and a taskbar of open panels along the bottom edge.
Six panels on one name, tiled. The strip at the top is the launcher; the left menu opens panels; the bar along the bottom is the taskbar of everything open.Captured from a live terminal on 2026-08-07 (5.19.6). The workspace chrome and the per-security menu are unchanged since; a chart panel has gained a 💾 save-settings control and the pane-height rule described below.

The one rule everything follows: a panel pins its security

A panel is opened against a security and a lens, and neither changes for its life. "AAPL · Chart" stays AAPL forever. Changing the lens means opening another panel; changing the security means opening another panel. There is no gesture and no agent call that retargets a live panel at a different name.

That single decision is what makes cross-security comparison possible everywhere — three Model of Models panels on three names, read against each other, on one screen.

Three Model of Models panels open simultaneously on AAPL, MSFT and NVDA, each reaching a different valuation verdict, arranged across the workspace.
The same lens on three names at once. Each panel is bound to its security, so nothing you do to one changes what another is showing.

The one exception: a panel whose security no longer resolves shows a refusal and offers a picker so you can deliberately retarget a panel that is showing nothing.

Where things are now

Part of the screenWhat it does
The top stripA launcher, not a mode switch. Each entry is a security context. Clicking one chooses which security the left menu will open next — it does not hide, show, close or open any panel. The strip is the union of contexts you added and the securities of every open panel. Alongside them sit the pinned tabs: 🔔 Alerts, 📁 Portfolios, 🧬 Cohorts, 🌐 Market, 🛍 Marketplace, 🌎 Macro-economic and 🛠 Administration.
The left menu (rail)Opens panels. Clicking a rail item opens a new floating panel on the current context. It marks exactly one item as focused — the view the focused panel is showing — and puts a quiet count beside any view that has open panels for this security.
The workspace barSits directly above the panel area (not across the rail) and wraps rather than scrolling. Carries ⊞ Tile · ⊟ Cascade · ▤ Stack · ⊙ Gather · ▬ All · ▣ Restore all, the panel and reduced counts, and any workspace notices.
The taskbarAlong the bottom of the workspace: every open panel, in open order, not just the minimized ones. Click to summon; middle-click or the hover to close; right-click for the panel menu.
A panel's title barThe drag handle, carrying duplicate, minimize, maximize ( when it is maximized) and close, plus a menu with duplicate, close others, bring to front and send to back. Double-click maximizes and restores.

Opening, duplicating and closing

Minimizing, arranging and never losing a panel

If a panel seems stuck: it is probably maximized. A maximized panel is pinned to the viewport, so its record can move while the screen does not — that is what a "snap back" used to be. Maximized panels now show it (an accent title bar and a control), and any drag or resize leaves maximized first, restoring under your cursor. Ask the agent to unmaximize_panel if you want to be certain of the end state; maximize_panel is a toggle and cannot promise one.

How tall a chart's panes are, and when the panel scrolls

A chart panel divides the height it is given between the price pane and each indicator pane. Two floors govern that, and both are derived from the drawing engine's own metrics rather than chosen: a pane must be able to show the engine's minimum number of price intervals with each interval two axis-label boxes tall (64px of plot, 90px for the one pane that draws the time axis), and its natural height is that same interval count at the spacing the engine is actually aiming for.

A layout saved before this rule existed reads as unknown, refuses to guess which of the two produced it, and resolves to fit.

Per-panel chart state

Two chart panels on the same name are ordinary, so chart state diverges fully per panel — including drawings. It is made safe by inherit-then-diverge: a new chart panel with no record of its own reads the existing symbol-scoped record and immediately writes a copy to its own key, so your existing drawings land on the first panel you open and on every later one, nothing is orphaned, and the original record is never deleted or written to by a panel. A standalone chart (a published chart URL) goes on reading exactly what it always read. A drawing that cannot be migrated is refused by name, never silently dropped.

The pinned surfaces

Administration, Portfolios, Alerts, Cohorts, Market, Marketplace and Macro-economic are full-pane covers, not panels. While one covers the workspace the taskbar stays visible and dimmed, and clicking an entry returns you to the workspace with that panel focused. The workspace is occluded, not destroyed: geometry, z-order, scroll positions and loaded data all survive and nothing is re-fetched.

What the workspace does not do (yet)

Stated plainly so you don't go looking: there are no named saved layouts on the server (what exists is export/import of the whole arrangement in one round trip, plus one-slot "restore previous arrangement" and "reopen closed panel" undos); no keyboard shortcut set beyond arrow-key nudging of a focused title bar; no screen-reader pass; no docking, snapping or tabbed panel groups; no tear-off OS windows; no cross-panel crosshair (the ⌖ Link control links the cells of one grid only); and no per-window workspaces — two browser tabs share one saved workspace and the last to save wins, which is disclosed on the workspace bar rather than prevented.

Driving it with the agent

You can navigate the terminal yourself, but the fastest way is to ask. The agent can read what's on screen and drive it.

You sayWhat happens
"Open a chart on AAPL and one on MSFT, side by side."Opens two chart panels and tiles them.
"Put the Model of Models for AAPL, MSFT and NVDA on screen at once."Three bound panels, arranged.
"I've lost the statements panel."summon_panel — restore, raise and bring into view, guaranteed.
"Tidy this up." / "Minimize everything but the chart."Gather / tile / minimize-all through the workspace verbs.
"What's open right now?"Reads the workspace state: every panel, its security, its view, its geometry, and whether it is on screen.
For agents and power users — a contract change worth knowing. active_symbol now means the security of the focused panel (what you are actually looking at), falling back to the strip's selection when nothing is focused. The strip's selection — the security the left menu will open next — is context_symbol. Both are always present, they may legitimately differ, and every payload carries active_symbol_means saying which of the two it currently is. Older verbs (set_view, select_security, add_tab, remove_tab…) still answer, each marking itself deprecated, stating what it actually did, and naming its successor.

Administration

The utility and engine surfaces live under one pinned tab — 🛠 Administration — behind a left-hand menu, grouped so the list reads rather than merely scanning. There are thirteen panels:

GroupPanels
MachineSettings (every API key) · Prompts (every text this installation sends to a model) · Chart preferences (what a new chart opens with) · Processes (the batches on this machine)
Event processing — what can be knownReach · Conditions
Event processing — what is watchedRules · Strategies
Event processing — what happenedEvents · Firings · Runs · Backtests
Machine (again)Support (guide, diagnostics, what's new and learning, served from the same knowledge base the agent uses)

The ordering is the reading order: what could be known, what can be asked, what is being watched, what is being held because of it — then what actually happened, what a rule has actually done, what a strategy did today, and what a replay of the past concluded.

Why "Machine" appears twice. The menu emits a heading when the group changes, in the registry's own order — so a heading that recurs is reported rather than tidied away. The tab's tooltip is computed from the same table by the same rule, which means it describes what you will actually see when you click it. It used to be a hand-written sentence naming three panels, written when there were three; it went unrepaired through eight more, on the only affordance that describes the tab.
Settings, Processes and Support were once top-level tabs; Thesis, Economy and Sectors were three more. Settings/Processes/Support are now Administration panels, and Thesis/Economy/Sectors — plus the release Calendar — are panels under 🌎 Macro-economic. Old links and old agent verbs migrate to the right panel rather than breaking.

Processes

Administration → Processes is the board for every batch this installation runs. Each row carries a status light with the reason for it, the cadence, when it last ran, when it is next due, and Start / Pause controls.

The Processes board listing the scheduled batches, each with a coloured status light, a plain-language reason, its cadence, last run and next due time, and Start or Pause controls — with several rows showing red for overdue or unscheduled jobs.
The process board, including its honest red rows. A job with no schedule anywhere on the machine reads red and says exactly that. Eleven processes are declared; the figure is the top of the board.

The status law

What runs

The registry is derived from the code rather than hand-maintained. It currently declares eleven processes: the surface batch, the screener enrichment batch, the daily repricing scan, the cohort refresh, the daily ledger, the monthly universe snapshot (the survivorship archive), the event-engine tick, the strategy loop, the monthly point-in-time snapshot, the legacy alert evaluation, and the economic-calendar pull.

An example of the house rule in action. The legacy alert evaluation ships on the board declared unscheduled — nothing on the machine runs it, and its row says so in a sentence rather than showing a stale green light. Naming a gap is the disclosure; closing it is a separate decision.

Green is not the only question

A process that ran cleanly over a near-empty input reports success — and success is not the question you were asking. That is not a false light (the run really did succeed) so it does not get a new colour; a colour would put a correct idempotent no-op in the same bucket as a job that covered 25 panels of 7,600. Instead every run also publishes a yield verdict from a closed vocabulary — full · partial · near_empty · empty_by_design · not_applicable — computed by the process itself.

Whenever the verdict is short, a not_this_process field is required: a yield disclosure that does not name whose shortfall it is converts a useful fact into a false accusation.

What a run cost

The heavy passes record what they spent — provider calls, calls per name, throttled seconds, HTTP 429s split by the population they came from, and cache hits — read off the request governor's own counter, so a pass that met a provider rate-storm can say so and a pass that could not measure its cost says that rather than reporting zeroes. A provider failure is classified by HTTP status and never by the provider's prose: a rate limit and a plan limit can arrive in the same English sentence, and a failure with no status is reported as unclassified rather than labelled from a guess.

Pausing is authoritative and journaled: a paused process reads red as "halted by owner" until you Resume. A red light is information, not a trigger — nothing on this board edits books, models or configuration.

Liveness is derived, never stored. A record that says a job is running is not evidence that it is: a dead process id is proof of death at second one, while a live one is never proof of life. So there is no state in which you can neither stop nor start a job — Start takes over a dead heartbeat and discloses it, and Stop against a dead loop retires the record and says plainly that nothing was stopped because nothing was running, rather than promising a stop no surviving process could honour.

One selector, two-tier disclosure

Two small things do most of the work of making the interface coherent.

One selector, one taxonomy

Every dialog in the system that asks you to choose something — a screen field, a cohort predicate, a chart indicator, a lens, a security, a transaction type, a rule's rank field — opens the same component. Same geometry, same search box, same keyboard (Esc closes, Enter commits, ↑/↓ walk the rows), same "no matches anywhere" sentence. Where a dialog deliberately behaves differently — single-select and Apply, versus a multi-stage basket and OK — the footer says which gesture it is using, so you are never guessing.

Everything selectable sits in one category taxonomy: fundamental · technical · model & attribution · macro & data · market & flows · identity & classification. Membership is derived from the underlying registries — screen fields, chart indicators, the series catalog, the pricing-model registry, model factors — rather than authored, so a field that becomes available becomes selectable, and an orphaned category fails the build.

Two-tier disclosure

The rule: if hiding it could cause a number on screen to be read wrongly, it stays on the row. If it only explains or attributes, it moves behind one icon.

Nothing is truncated behind the icon: the full text renders with no line clamp, and the popup is keyboard-dismissible with focus returned to the icon that opened it.

Charts & the workbench

The charting workbench is a full technical-analysis surface with its own engine — ~125 indicators, 39 drawing tools, all timeframes and extended hours, compare overlays, and event markers. Open one with a request, then refine it in words. Inside the terminal it is a panel; you can also open a standalone chart page.

The charting workbench showing candlesticks with moving averages, a lower oscillator pane, drawing tools in the left toolbar and sentiment-coloured event markers along the price series.
The workbench. Indicators, overlays, drawings and event markers are all addressable by the agent.Captured at 5.19.6. The legend row has since gained a 💾 save-settings control whose label states which preference is in force, and — when analyst rating changes are on — a ⛃ Firms… control at the right edge of the events row.
You sayWhat happens
"Chart AAPL daily with the 50 and 200-day moving averages, RSI, and MACD."Opens a chart panel with those overlays and oscillators.
"Add an anchored VWAP from the last earnings date and mark the gap."The agent adds the drawing/overlay to the panel you're looking at.
"Switch to weekly, log scale, and compare against MSFT and the SP500."Reconfigures timeframe/scale and adds normalized compare lines.
"What do the signals say?"Reads the live 12-signal technical battery + key levels off the chart.
Charts are served from a localhost URL; the API key stays server-side. A standalone file copy is also saved — that copy embeds the key, so treat it like a credential. Inside the workspace, an embedded chart holds no long-lived connection however many chart panels you open, so a screen full of charts does not starve the terminal of sockets.

Chart preferences — the defaults stop being a secret

Every new chart used to open with a 50 and 200-day moving average, RSI(14), MACD(12,26,9) and a volume underlay, and that set was written down nowhere you could reach — it was transcribed four times, in three spellings, across two documents, and the volume underlay appeared in none of the server-side lists at all, which is precisely why it could not be found. Two of the four disagreed about the MACD parameters, and agreed only for as long as the indicator registry happened to say 12/26/9.

WhereWhat it sets
Administration → Chart preferencesThe universe default — the indicators and parameters every new chart opens with, chosen through the same selector every other dialog uses, over a list derived from the chart's own indicator registry.
💾 on any chartSaves that chart's settings for that (security, interval) pair. The control's label states what is currently in force; Alt-click clears it; every saved pair is listed in the panel with a Clear beside it.

A record captures indicators and parameters, overlays, comparisons, style, log/linear, the volume underlay, event layers and the window. Drawings are deliberately not included — a drawing is an annotation rather than a configuration, and it already carries its own per-panel inherit-then-diverge contract.

The key is (security, interval), and there is no fall-back between intervals. AAPL daily and AAPL weekly are different records, because a 14-period RSI is a different study on each — so the interval stops being a saved value and becomes part of the address. Resolution runs explicit caller → saved (security, interval) → universe default, and the payload names which rung answered. Where a chart falls through to the universe default, the legend says so in words and names the intervals that do carry a record, because "where did my settings go" has to be answerable from the screen.

A preference is installation state, so it deliberately does not travel in an exported layout — and the export says so rather than omitting it silently. Importing a layout must not overwrite a preference you set for every other chart.
Administration, Chart preferences: a table of the universe-default indicators with their parameters and a column stating where each parameter came from, above a table of saved per-security, per-interval records each with a Clear control.
The universe default, with a column stating where each parameter came from — chosen in this preference, or inherited from the chart's own indicator registry. Beneath it, every saved (security, interval) record with a Clear beside it.

Analyst rating changes, sliced by firm

The rating layer used to draw all of them or none of them, and on a well-covered name that is a wall — a single measured grade stream ran to 1,784 rows. Turn the layer on and a picker offers All or a list of firms, each with its change count over the chart's currently loaded bar span, so the number beside a firm always describes markers you could actually see.

Nothing is scored. A per-firm accuracy number needs a horizon, a benchmark and a definition of a correct call — and on a well-covered name reiterations are the majority of the rows. Rather than invent one, the payload carries the four decisions such a number would require, so a later version inherits them rather than reinventing them.

Chart grids

A four-up chart grid panel: four securities charted on a shared date window with a linked crosshair, each cell carrying its own indicators.
A 4-up grid with synchronized ranges. The grid is one panel; its identity is the ordered set of securities plus the layout.

Pan or zoom one cell and the others follow. The crosshair relay resolves to its owning panel, so it never leaks into a different grid. Ask for "a 4-up grid of AAPL, MSFT, JNJ and KO" and re-asking focuses the grid you already have; Alt-click the menu item (or ask for a duplicate) to get a second one.

Overlaying macro, commodity, financial & filing series

The chart isn't limited to price. The Indicators dialog overlays any external time series — economic (FRED), commodity/energy prices, per-company financial-statement line items, or SEC EDGAR XBRL facts — directly on the chart so you can compare and test predictive relationships. Each overlay renders either on its own stacked right-hand axis (so different units line up honestly) or in a separate pane, your choice per series, and can be shifted forward or back in bars to line a leading indicator up against price. Ask the agent to "find the best lead/lag" and it runs a cross-correlation (analyze_lead_lag) to tell you which lag aligns them best and how strongly.

You sayWhat it does
"Overlay the 10-year Treasury yield and WTI crude on this chart."Adds each as its own right-hand axis (refs like fred:DGS10, commodity:CLUSD).
"Does copper lead this stock? Shift it to the best lag."Runs the lead/lag cross-correlation and applies the shift that best aligns them.
"Plot AAPL quarterly revenue and its EDGAR net income under the price."Adds fmp:income:AAPL:revenue:quarter and edgar:AAPL:NetIncomeLoss as panes.
Overlays align to the price dates by as-of (forward-fill), so a monthly or quarterly series steps correctly against daily bars. Lead/lag shifts and correlations are analytical aids — verify the economic logic; correlation is not causation.

Relative views: Δ$ and Δ%

A macro series on the chart is rarely expected to equal the price — the relationship is the object. Every overlay's row in the picker offers three renderings: value (as fetched), Δ$ (price − series), and Δ% ((price − series)/|series|). A model line flipped to Δ$ reads as the premium itself, through time.

Replay without lookahead, and rule backtesting

Replay walks a chart forward bar by bar from any start date. The no-lookahead guarantee is structural, not a promise: the chart is windowed to the cursor, so indicators, the 12-signal battery, auto-detections and event markers all recompute from bars at or before the cursor and simply cannot see the future. Fundamentals get the honest treatment most tools skip — analyst targets and DCF lines are as-known-now, so during replay they are suppressed and labelled unless one of your point-in-time snapshots covers the cursor date, in which case the values you actually observed then are drawn and marked as-known-then.

Rule backtesting takes entry/exit rules in the series DSL (or describe them and the agent writes them) and returns the trade list, equity curve, hit rate, average win/loss, profit factor, exposure, max drawdown and a buy-and-hold comparison, with costs in basis points per side. Every rule backtest reserves the most recent ~25% of bars as a holdout and scores both windows — the split happens inside the engine, so nobody can request the flattering half alone.

A single-symbol rule backtest is a hypothesis check, not strategy validation. Costs are modelled per side; slippage, borrow, taxes and survivorship are not — and rules chosen by looking at the same sample are an upper bound. Note a level rule (close > 50) is a state that re-enters after each exit; use crossover/crossunder for edge-triggered entries. For portfolio-level testing, see the backtester.

Indicators you define yourself

Beyond the built-in library you can define your own indicator from a formula and use it exactly like a built-in. Ask for it in words — "define an indicator called vol_ratio that is the 10-day average volume over the 50-day average volume" — and it appears in the indicator library under Custom, with the same add / edit-parameters / hide / remove controls as everything else. Definitions persist.

Formulas are written over the bar fields (open, high, low, close, volume) with windowed functions — sma, ema, wma, highest, lowest, stdev, sum, slope, ref, crossover, crossunder, barssince — plus any parameters you declare, which then become editable inputs on the chart.

The formula language has no loops, no assignment and no access to anything outside the bars — it is evaluated on the server, never eval-ed in your browser, and the chart sends bars it already holds, so a custom indicator costs no extra API calls. The same expressions drive rule backtesting, so an idea you can plot is an idea you can test.

Auto-detected trendlines, zones & patterns

The 🔎 Detect menu runs pattern recognition locally on the bars already loaded — no extra API calls. Toggle Trendlines (fitted through pivots, scored by touches, recency and violations), S/R zones (volume-weighted pivot clusters), Candlestick patterns (~40 recognizers), and Geometric patterns (double tops/bottoms, head & shoulders, triangles, channels), with a minimum-confidence slider. Detections render as ordinary chart drawings but tagged auto — dashed and tinted by direction — so they never mix with your own drawings and can be cleared separately.

Crucially they are structured objects, not pixels: each carries chart coordinates, a confidence score, a direction, a status (forming / confirmed / invalidated) and a plain-language narrative, so the agent can read exactly what you are looking at and act on it.

Detections are structured estimates from price geometry, not predictions — forming means the confirmation trigger (neckline, breakout) has not happened yet. Confirm against context before acting.

Events & sentiment on the chart

The 📅 Events menu covers every source: earnings (beat/miss), dividends, splits, insider buys/sells, SEC filings, earnings-call transcripts, institutional 13F ownership flow, analyst rating changes, quarterly financial releases, news, your own research notes, and events you have associated with a series — which draw wherever that series is overlaid, on whatever security's chart that happens to be. Each is a separate toggle, and every marker is sentiment-colored, with busy days collapsed into a single cluster. Turn on Show events pane for a lane charting net event sentiment per bar.

Hover any marker to read it; click it and the terminal opens the view that owns the source — filings to SEC Filings, ratings to Analyst Ratings, insider to Insider Trades, an event you entered yourself to its calendar drill-down — with the dated row highlighted. If the security has an operating model, every event also carries a model link naming which driver it bears on, drawn only through edges your model really has.

Every marker now says where it came from, first, above the label. Eleven of the twelve sources had no provenance at all until recently — a dated row that cannot say where it came from must not sit on a chart where a reader cannot tell it from one that can. Two origins are named rather than flattened: eleven sources are a publisher stating something, and a research note is you stating it — the same origin an event you add to the calendar carries, because it is the same kind of claim. A row arriving with no provenance renders "source not stated by the producer" rather than as though it had one; silence there is indistinguishable from a publisher.
Chart event sentiment is a heuristic read per event type — a fast orientation aid, not a verdict. Read the underlying event before acting on it. Note the deliberate contrast with the typed event log, which withholds sentiment entirely: the chart's colouring is presentation, not a fact recorded about the event.

Screening stocks

The screener does cross-sectional factor ranking (percentile or winsorized z-score, sector-neutral), an expression language for custom filters and derived fields, quality/health scores (Piotroski, Altman), 12-1 momentum, 13F crowding, and point-in-time-honest factor backtests — over a vocabulary of 520 declared fields, including the enrichment families and 148 technical readings derived from the chart's own indicator taxonomy.

You sayWhat it does
"Screen large-cap tech under 20× earnings, ranked by momentum, sector-neutral."Builds and runs the screen, returns a ranked table.
"Add a filter: free-cash-flow yield > 4% and Piotroski ≥ 7."Uses the expression DSL to refine the universe.
"What fields can I screen on?"list_screen_fields — the catalog, plus the register of what is deferred and what is not planned.
"Save this as 'Quality Compounders' and tell me who entered or left since last run."Persists the screen and reports membership deltas over time.

What the vocabulary says about itself

Modernized G & D

Graham's defensive criteria, retranslated instrument by instrument. The diagnosis behind it is not that the criteria are too strict but that the instruments now adversely select: a price-to-book gate in an intangible economy is close to a filter for asset-heavy, low-return, structurally declining businesses, because those are the only firms whose capital base still shows up on the balance sheet at something near economic size.

Ask the agent to list your screens and cohorts to find them; each shipped cohort points at its screen by name, so membership re-resolves from the rule rather than freezing a list.

Read an empty result correctly. On its first live run every variant returned zero members, and every shipped description says so: the count is a statement about prices at that moment rather than about the instruments, the hurdle is demonstrably attainable, and nothing was tuned to produce a list. Before reading an empty result as a finding about prices, check surface coverage — on a fresh install the value side reads a recorded surface day, and if none has been recorded there is nothing to be cheap against.

Cohorts

The Cohorts tab listing peer groups with their kind, member counts, premium and coverage headers, and freshness disclosures on each row.
Cohorts are lenses, never conclusions — the same groups drive comps, the price decomposition, the heat-map slice, the repricing report, the alert feed and the calendar.

The 🧬 Cohorts tab maintains the peer groups every valuation is measured against — sector, size, index, factor screens, and your own lists, with equal / cap / explicit weighting. A cohort detail view carries premium and coverage headers, a modeling-gaps read, and view on surface. A thesis can materialize its screen into a living cohort whose membership re-evaluates as fundamentals change.

A portfolio-linked cohort derives its membership at read time from the linked portfolio's open positions, so there is no stored list to fall out of sync — the resolution is the sync. These cohorts cannot be edited or deleted directly; the refusal names the derivation.

Cohorts built from live listings are survivorship-biased and say so in every response. A monthly membership snapshot accrues toward a genuinely point-in-time universe, forward from first use — which is exactly what the Reach panel reports on and what a cohort-scoped historical backtest refuses without.

The pricing engine & the Model of Models

More than twenty valuation frames behind one interface — every methodology the industry uses, presented in the industry's own terms — plus a Model of Models that estimates the mixture of frames the market is actually pricing each name on. Two panels carry it: 💠 Pricing Models (every model side by side) and 📐 Model of Models (the precision-weighted blend, with the modeling-gaps strip).

The Model of Models panel: each valuation frame listed with its fair value, its walk-forward weight and its distance from price, above the precision-weighted composite, with declined frames stating their reasons.
The weights are the reading. A frame that cannot value this name declines with its reason rather than contributing a silent zero.

Diagram of the pricing engine: twenty valuation models feeding correlation weights, producing the model of models composite, plotted on the chart

You sayWhat it does
"Compare every pricing model on AAPL."Runs the whole basket — intrinsic (DCF, FCFE, EPV, DDM, residual income, EVA), relative multiples at the name's own trailing median, asset models, Monte Carlo, scenarios, SOTP — each with its fair value vs price; models that decline say why.
"What frame is the market pricing TSLA on?"The Model of Models: each model earns a walk-forward correlation weight against the name's own price (strictly as-of, no lookahead). The weights are the reading — and their year-over-year drift shows the rotation.
"Chart the DCF and EV/EBITDA lines against price."Every model has a point-in-time series (filed statements only) — add from the chart picker's Pricing engine section, or by ref: pricing:AAPL:intrinsic.dcf_fcff.
"Replot the composite under January's weighting."Weight vectors are recorded per stock per date; pricing:SYM:composite.momw:2026-01-15 holds that day's recorded vector fixed across the whole series.
"Show the rotation itself."Per-model weight series — pricing:SYM:weight.relative.pe and friends — chart each model's weight through time in percent.

Valuation structure: what kind of thing is this company?

It is naive to treat every frame as an equal peer, and it is wrong to blend an add-on book as a peer of a whole-company frame when it is an increment. So each name carries a declared valuation structure, exactly one in force at a time, dated and attributed:

Classification is evidence-proposed and owner-overridable, and carries the condition under which it should be re-proposed. In the shipped default the structural composite is computed and disclosed beside the incumbent composite rather than replacing the published headline — disclosure precedes adjustment.

Judgment models & strategic-initiative books

Some of what the market prices does not exist in filed statements — capital programs and pre-revenue ventures like robotaxi fleets, humanoid robots, AI infrastructure. The 🎯 Strategic Initiatives panel maintains full SI books: per-initiative dated cash-flow lattices with capital schedules and ramps, milestone gates with expected resolution dates, a risk register, an append-only event journal, immutable snapshots, maturation fade, and calibration scoring — plus a belief dial that inverts the market price into the implied success odds, recorded as a daily series. Editor views include the tornado, waterfall, and probability cone. Alongside the books, simpler judgment models take the inputs only an analyst can supply:

ModelYou maintain
Venture bookPer-venture TAM, achievable share, margin at scale, probability, years, committed capital. Valued as probability-weighted PV net of committed capex — a venture whose spend exceeds its prize shows a negative EV, as information.
Real optionsOpportunity value S, investment cost K, horizon T, volatility σ (defaults to the name's own).
Platform economicsUsers, ARPU, contribution margin, capitalization multiple.
Qualitative scorecardFour 0–100 analyst scores; the engine computes the weighting and invents no score of its own.
SOTP tiltsPer-segment EV/Sales multiples over the disclosed segments.

Saved inputs persist per symbol with an edit history and activate the model everywhere. Every stored number is labelled analyst judgment. The pairing that matters: chart the composite in Δ$ to measure the premium the market is granting for the unseen businesses — then write your own venture book and compare.

Company operating models

Give any company a driver-based operating model calibrated from its filings and consensus, then turn any event into an instant EPS and fair-value delta. An elasticity table shows exactly which drivers move the stock. The operating model is the machinery beneath the pricing engine's intrinsic models — its drivers feed the DCF, FCFE, EPV and Monte Carlo lines — and in the terminal its panels are superseded by the Model of Models view; the model itself remains fully operable through the agent.

You sayWhat it does
"Build an operating model for CAT."Calibrates revenue → EBIT → EPS → FCF, plus exit-multiple + DCF fair value.
"Shock gross margin down 2 points and raise the tax rate to 21%."Re-runs the model and reports the EPS and fair-value impact.
"Which drivers matter most?"Returns the elasticity table (sensitivity of fair value to each driver).

The influence map — the outside world, wired in

A model isn't just its internal drivers. The influence map wires external economic, policy, supply-chain and consumer/customer factors to the internal drivers they move by a signed influence edge with a stated, auditable rationale. It renders left-to-right — factors → drivers → EPS & fair value — with green edges that raise a driver and red that lower it, thickness scaled to sensitivity. Shock a factor and Open Trading Surface propagates it through the edges into the drivers and reuses the valuation engine, returning the full factor → driver → outcome trace.

You sayWhat it does
"Show CAT's influence map."Opens the factor → driver → value diagram (seeds a sensible default factor set the first time).
"Raise tariffs 10 points on CAT — what happens to fair value?"Propagates the tariff shock into gross margin and re-runs the model, with the trace.
"Calibrate CAT's factors from data."Pulls each factor's current level from FRED and, on request, fits sensitivities by regression against the company's own history.
Influence edges start as structured estimates. A fitted coefficient only replaces the default when the fit is trustworthy (enough aligned years, real variation, a real correlation); otherwise it is attached to the edge as auditable evidence and the prior is kept. The regression cross-check can mislead under collinearity or short history — treat disagreement as a prompt to re-examine the edge, not proof either way.

Forecast attribution & weighting

Any external series can be a model dependency, and the model will tell you what actually carries the forecast. Run forecast & attribution projects every series-backed factor forward, propagates each through its weighted influence edges, and decomposes the multi-year fair-value / EPS forecast into per-factor contributions with computed weight shares (ranked bars, plus a reconciliation residual for interaction effects). Each factor carries a tunable weight (1 = as-calibrated) to amplify, damp, or mute a dependency without touching the calibration.

Attribution is a decision aid, not a forecast guarantee: contributions are marginal (one factor at a time) so they won't sum exactly to the total — the residual is reported.

Going deep: segments, quarters, and the full three statements

Segment-level drivers. Open Trading Surface assembles each company's disclosed segment revenue history, ties every period to consolidated revenue (a residual becomes an explicit unallocated segment; an overshoot is flagged suspect, never scaled away), and detects re-segmentations as structural breaks. Drivers may then be scopedrevenue_growth_pct@cloud — and the projection blends growth by the segment mix, with mix shift as an output you can chart.

Quarterly frequency. Annual fits cap near 15 points; quarterly gives ~60, which honestly earns the A evidence grade that annual data structurally cannot. Seasonality is treated as structure, never signal: all quarterly fits use year-over-year same-quarter changes. Each dated model snapshot becomes a next-quarter EPS checkpoint scored against the print within weeks.

Three-statement articulation. Setting a payout ratio, a buyback share of FCF or a target leverage switches a model into articulated mode: earnings roll into equity, capex − D&A into fixed assets, FCF less dividends and buybacks into cash, and interest is recomputed from average net debt. Buybacks retire real dollars. The tie — assets = liabilities + equity — is enforced in every projected year: a model that cannot tie refuses to run and names the gap.

Honesty rules carry through: segment coverage is uneven below large caps and absence renders as absence; the quarterly EPS checkpoint is a derived read of the annual model, labelled as such; articulated mode refuses to run without an opening equity rather than fabricate a balance sheet.

Pro forma statements & the substrate

A financial statements panel showing an income statement across several annual periods with derived margins and growth rates in the terminal's tabular style.
Filed history and projected periods on one grid. The statements are where every layer of the stack actually converges.

Beneath the valuation frames sits a driver-based three-statement pro forma: normalized historicals on the left, forecast periods to the right, statements that articulate, and valuation computed off the projection. Its layers, in order:

The assembled object is the substrate: a grid of quarter labels (filed and projected on one calendar), the statements themselves, a stamp of the layers that produced them (economy state, industry chain, driver fits, initiative book generation, financing policy), and an as-of date. It is point-in-time by construction: a substrate struck at a past date contains only filings whose filing date precedes it, and where the economy container has no snapshot at that date the current one is used and stamped as anachronistic rather than silently substituted.

What this means when you read a valuation today. The substrate object exists and is readable (get_substrate), and it is what the direction of the system is toward: models as independent lenses over one statement set — each free to disagree about what the statements are worth, none free to disagree about what they say. That migration is staged, and in the shipped default the precision-weighted composite still carries the headline. As with every layer here, the change will arrive as a disclosed, dated methodology break rather than a silent one.

Simulation is the intended run mode: one draw from the joint driver distribution produces one articulated statement path that all readouts then read, with gates becoming Bernoulli draws rather than probability multipliers. Quarterly periodicity out to 24 quarters, a duration set, versioned pro forma snapshots and event-to-line revision are specified and not yet built.

The economy & industry containers

Under 🌎 Macro-economic sit three panels: Thesis, Economy and Sectors — a belief about the world, the economy that prices it, and the industries that carry it.

The economy is a container, not a forecaster

It does not forecast, does not opine, and computes nothing a valuation could not compute for itself given the values. Its job is that there is one place the values live. The admission test for a variable is narrow: does it change a number in a valuation? — the real growth rate, inflation, the policy rate, the ten-year risk-free, the BBB credit spread, the equity risk premium, potential growth, the effective tax rate.

Variables are series on the pro forma's own period grid, not a row of current values, so inheritance downstream is a lookup rather than an interpolation, and a grid mismatch is a refusal rather than a coercion. Scenarios are a complete alternate filling of the container with probabilities that must sum to one — a scenario is not a stress test.

Why this exists. Every valuation already contained an implicit economy model — an equity risk premium compiled into a cost-of-equity block, a risk-free fetched per name, a terminal growth rate defaulting to a bare number — one that was never stated, never versioned, and could not be changed without editing code. Making it explicit is the point. Adoption is staged per input and disclosure first: consumers compute both values side by side and report the pair, the delta, and the resulting valuation delta per name, with the published headline byte-identical until the switch is thrown — and then it is a dated methodology break, because the number changed when a scattered implicit economy became one explicit one, not when the world did.

The industry rung

Sectors is the same shape one rung down. Classification is the provider's, and the legal key sets are derived from the provider rather than curated, because a curated list rots. Resolution runs coarse to fine, the finer winning — sector container, then industry container, then company override — with the whole chain disclosed. Variables are industry-native quantities that do not exist at the economy level: industry revenue growth, operating margin, capex intensity, beta, credit spread.

The governing rule: a fitted driver is a measurement; an industry variable is a prior; a measurement is never silently displaced by a prior — not with every toggle on. The visible payoff is the case that previously got nothing: a thin-history name whose driver fit refuses now shows the industry prior as a disclosed, sourced, provenance-stamped option — refusal plus an honest alternative, never a silent default.

On top of the containers sits a dependency layer: one dependency is one claim — this driver, in this industry, moves with that variable, by this much — with elasticities as time series on the pro forma grid (a scalar is accepted as shorthand for a flat series and stamped as such, because "flat" is a stated shape, never a hidden assumption). Three layers: a common template, per-industry templates written the way sector analysts model them with their basis stated, and your own overrides, which require a reason. Templates never claim measurement — a payload calls a template a template — and dependencies condition drawn scenario worlds without shifting the deterministic base projection at any setting.

The economic calendar

🌎 Macro-economic → Calendar is the schedule of economic releases — what is expected, what printed, and what the surprise was — with a route from any release to the chart of the series it publishes. It needs a free FRED key; without one it says so rather than showing an empty grid.

Four zooms, and each one is a real grid

ZoomWhat you see
AnnualTwelve month mini-grids, every day a cell, shaded by intensity. A year is roughly eleven thousand occurrences, so this view shows shape and says so on its face. Every month is present including the empty ones — a missing March would shift every later month into the wrong slot.
MonthlyA real seven-column grid. Each cell carries the full count plus ranked chips and a real +N more.
WeeklySeven day columns, Sunday-anchored. It used to be today − 3 → today + 3, which is seven days and not a week — and a zoom whose window is not aligned to its unit cannot be a grid.
DailyEverything for that day, in full.

Every shade is a ratio against a published denominator — the busiest day's count, printed in the legend. An intensity with no denominator is a palette.

The economic calendar at the monthly zoom: a seven-column grid of August 2026, each day cell carrying its occurrence count, ranked release chips and a plus-N-more link, shaded by intensity, with the link-rate and consensus-archive disclosures printed above the grid.
The monthly zoom — a real seven-column grid, each cell carrying the full count, ranked chips and a real +N more. Above it, in words: the link rate, the sentence saying it is a floor and not a measurement because the second matching lane is only partly warmed, and the size of the point-in-time consensus archive.

The ‹ › buttons and the agent's own move verb read one derived navigation block rather than each computing a date, because two implementations of one movement put the first month-boundary defect in exactly one of them. A direction with no target refuses by name — annual has no coarser zoom, and the refusal names the range bound that is the reason rather than pretending it is an omission.

Consensus is recorded before the print, not after

The pull runs twice on weekdays. That is not redundancy: an evening-only pull structurally cannot record a consensus before an 08:30 print, so it would be a different and broken process. What was expected is written to an append-only point-in-time archive as it was believed at the time, and the surprise is computed against that rather than against a figure revised afterwards.

Jump to the chart — of a security you choose

Pick a release, pick one of its series, and the terminal asks which security's chart to put it on. That question is the feature. The obvious behaviour — open a bare chart of the series — answers show me this series when the question was show me this series on something, and produces a chart with no price beside it and no relation to any name you were working on.

The link rate is published, and it is not high

Each dated occurrence carries a consensus only if a row from the estimates provider can be matched to it, and that match rate is 22 of 100 on a measured window — a quarter of what the design assumed before it was measured.

It is low for a structural reason no threshold fixes. The two providers name different levels of one taxonomy: one names the series ("Initial Jobless Claims"), the other names the release ("Unemployment Insurance Weekly Claims Report"). A second matching lane over FRED's own series titles took the rate from 12% to 22% at unchanged precision — and it is still derived rather than read from a table. One case is provably unreachable and is recorded rather than tuned around: "Inflation Rate YoY" against the Consumer Price Index, none of whose series titles contains the word inflation. The honest answer there is a refusal, and asserting the link yourself is the escape hatch.

Two things follow that are worth knowing before you read the number:

Your own calendar events

Forward dates are either known from a reliable source or left out — there is no periodicity-inferred "next" date anywhere in this system, and that is a rule rather than an omission. The most tempting derivation is the one that is so nearly always right, which is exactly what makes a wrong one unfindable. What replaces it is you: add an event with a date and a title, and associate it with a FRED series, an EDGAR series, a security, or nothing.

They appear at all four zooms, on the chart, and in the unified Alerts calendar — under the same event ids, so the same thing on two surfaces is the same thing.

Series, overlays & knowability

Everything periodic in this system is one observable series: a dated observation stream, its metadata, and its release schedule — with three renderings. Past observations plot. Dated occurrences become markers on a price axis. Scheduled releases become calendar rows. The hybrid is the proof the abstraction is right: history plots and the schedule marks, on one chart, with today as the boundary between data and schedule.

One identity, and a picker that stopped having a list

A reference like fred:DGS10, edgar:AAPL:NetIncomeLoss, fmp:income:AAPL:revenue:quarter or pricing:AAPL:intrinsic.dcf_fcff resolves through one identity module with one normalization rule per source, one constructor and one parser. Before that, twelve places in the code built references by string concatenation — twelve independent untested implementations of the same identity.

The Indicators picker no longer holds a hand-written catalog. Each consumer publishes the series it declares and the picker derives the union, in three tiers:

TierHow manyHow you reach it
A — declared28Listed outright. This is the union of what seven modules in the server actually use.
B — reachable~2,360 on a warm installationBrowse FRED's releases as nav nodes, alphabetical with each node's recorded occurrence count beside it — a nav tree that reorders itself between sessions is a nav tree you cannot learn. Series arrive on expansion.
C — searchable~800,000Search only. Results carry FRED's own popularity ranking, printed, and the heading claims only that.
Nineteen of the twenty-eight declared series used to be invisible — among them the potential-growth series that caps the terminal value of every valuation in the system. That count is now a derivation rather than a hand transcription, and a test fails the build if any series named anywhere in the server cannot be found in the picker.

Tier B is a property of what this machine has pulled, so an empty branch would read as "FRED has no releases." It never does: three cold states are kept apart — never pulled (naming the pull process and the two doors that still work: search, and typing a reference directly), warm-up short (printing what it has and saying the count is a floor), and genuinely empty (a release with no chartable series, which is the world being empty and a different fact from both).

Search quality is not measured, and the surface says so rather than printing a number. The method is fixed in shipped code — twenty declared queries — and the surface publishes either a measurement with its date and per-query detail or an explicit refusal naming the script that would produce one. A precision figure typed into a source file is a claim with no evidence behind it, and a build fails on a record that carries a precision without a measurement behind it.

The two dates, and the dated methodology break

Every observation has two dates: date — the period the number is about — and knowable_at — when it became knowable to you. A plot is dated by one and an inference by the other, and that is why one series cannot carry a single date field and serve both.

Filed EDGAR figures used to carry only the period end. A quarter ending 30 December was therefore treated as knowable on 30 December, weeks or months before the filing that published it existed — while the statements branch a few lines away did the opposite and disclosed that it did. That is fixed, and it moved published numbers on purpose:

What movedBy how much
As-of cells on a typical net-income series196 of 204 move (96.1%), median lag 578 days
An instant-in-time concept like total assets47.5% move, median lag 33 days — the clean case
Corpus-widemedian 496 days, 95th percentile 849, maximum 1,518
One worked valueAAPL net income as of 2019-01-31 was $19.965bn and is now $8.717bn
Two limits stated plainly. Filed statements from the market-data provider still date their points by the filing date while EDGAR dates its points by the period end, so two fundamental overlays can sit at different x positions for the same fact — each says which basis it used, and harmonising them is a further dated break that has not been taken. And an EDGAR series can serve several period durations under one line: a measured example carries 40 quarters, 19 full years, 8 nine-month year-to-dates and 7 half-years drawn as one line. Every read now computes and publishes that composition on the series in front of you, because the honest fix needs a grammar for period duration and would change every existing EDGAR series.

Perspectives — the three lenses

🔭 Perspectives is a per-security, model-written assessment. Configure a key for one AI vendor (see Configuring API keys) and the panel collects what the server already holds for that name, passes what is necessary to the model, and asks for an assessment grounded in classical value discipline.

The report comes back as prose in structured sections — business and earning power; balance sheet; valuation against substantiated value; temperament; risks — ending with an overall defensive / enterprising / speculative characterization, a six-axis characterization chart (financial strength, earnings stability, earnings power, growth record, dividend record, margin of safety), and the house disclaimer.

Three lenses, differing in one thing

It ships with three default lenses — bear, neutral and bull. They are not three tones of voice and they are not three conclusions. A lens is a rule about where the burden of proof sits: under the bear lens the favourable reading of a figure must be earned from the record while the unfavourable reading stands until the record answers it. Everything else — the instruments, the accounting gates, the refusal to credit forecasts, the required disclosure of what could not be checked — is identical across all three, because those are the discipline and the lens is only the burden.

The prompt library — Administration → Prompts

Every text this installation sends to a model lives in one place, with the sections derived from the store's own scopes rather than a hand-written map:

ScopeWhat it is
Lens defaultsThe three shipped texts. Amend, reset, and read the full version history.
Per-security overrideReplaces a lens for one name. The panel marks it ✎ overridden here.
Per-security libraryNamed prompts you write for one name. Structurally distinct from an override — an override carries the global chain it overrides, a library entry carries none, and a write that wears the other's clothes is refused.
PromotionTurns a security prompt into a shared one. This makes scope a property of a version, never of a chain — otherwise promoting a prompt in September would rewrite what every August report claims about itself.
DiscoveryThe thesis-discovery prompt, which had four agent verbs, no route and no surface anywhere until this panel existed.

Adopting a per-security override as the global default is offered, and it is disclosed as a dated methodology break rather than presented as a save.

Administration, Prompts: the six prompt scopes listed in order, each shipped lens default shown with its version, character count and creation date and Edit, Versions and Reset controls, with the empty per-security scopes explaining where they are created from.
Every text this installation sends to a model, in one place. A save never edits a version — it appends v(n+1) and makes it the head, because every stored report pins the version number that produced it. The two empty scopes say where they are created from rather than rendering as blank.

How it behaves

What the bundle says about itself

The facts sent to the model come with a manifest: what was included, what refused, what was dropped for budget, and — the load-bearing half — what the assembler was capable of at that build. A report written before a block existed must not read as though the model saw that block and dismissed it.

Every refusal reports the request that was actually sent — the exact body with credentials redacted and the prompt and bundle replaced by their lengths, beside the vendor's own token counts. A refusal that reports a number the request did not carry cannot tell the bound was never sent from the bound was sent and ignored, and that ambiguity cost two releases to resolve. Relatedly, the shape of the reasoning parameter is negotiated by the request rather than guessed from a model name: a rejection naming the current shape's own parameter advances to the next, and the answer is learned per vendor and model. A table of model ids and their capabilities would rot on the vendor's release schedule rather than on ours.
Older reports display honestly rather than being backfilled: reports written before the prompt was versioned say "prompt: unversioned", reports written before the characterization existed say "no characterization recorded", and reports written before the vendor was recorded say so and show the old provider field. Asserting a record was written under a vendor id that did not exist when it was written would be a fabrication.
Vendor errors surface verbatim — a bad credential shows the vendor's own message, an endpoint with no model-list route says so and invites a typed model id, and no credential shows one sentence naming both places to configure one plus the guarantee that nothing is sent anywhere. There is no simultaneous-vendor mode, no automatic failover, no scheduled auto-rerun, no ensemble, and no cost estimation: the panel says a run spends credits and does not pretend to price them.
A perspective is a model-generated interpretation of recorded data. It characterizes; it never instructs. It is not investment advice and it may be wrong.

The thesis engine & discovery

Turn a view into a maintained, probability-weighted thesis with an evidence ledger — and let an AI agent propose its own from this system's own planes. As evidence arrives, the probability moves and the change is logged. As theses resolve, Open Trading Surface scores its own calibration (Brier score, skill vs base rate). The Thesis panel lives under 🌎 Macro-economic, with every thesis, the + New thesis affordance and the 🌍 Worldview nested under it.

You sayWhat it does
"Create a thesis: US energy-security capex supercycle; sectors Energy and Utilities overweight."Records the thesis with an initial probability and tilts.
"Nat-gas guidance was cut — log that as evidence against, and reassess the probability."Appends to the evidence ledger and updates the probability with a reason.
"Run thesis discovery."Assembles the bundle, stages the run, and writes candidates into the review queue.
"Compile this thesis into a screen and build a paper model portfolio."Proposes a screen for review, then materializes a cohort from the accepted one.
"How well-calibrated have my probabilities been?"Reports calibration once enough theses have resolved (else says so honestly).

The journey: four stages, one panel

The panel used to be one long page that grew as steps were taken. It is now a staged progression, and the stage is derived from what the record carries and never stored — a stored stage is a fact duplicated into a field that can disagree with it.

StageWhat it means
1 · CandidateA proposal in the review queue — model-written or your own — to be read, edited, rejected or converted.
2 · InventoryA thesis you have decided to manage. This is a legitimate resting state: a thesis may sit here indefinitely with no screen, and the surface may not nag about it — no severity, no warning, no completion fraction.
3 · ScreenA screen candidate proposed, reviewed and accepted. See below.
4 · CohortThe accepted screen materialized into a living cohort whose membership re-evaluates as fundamentals change. This is the terminus, declared as one — a model portfolio still works and is simply not a stage.

The current stage renders open, earlier ones collapsed and reachable, later ones named and not open, with one advance affordance on the row. Selecting a thesis in the left menu opens it at its current stage; the 🧪 Candidates queue stays its own destination, because reviewing eight candidates in a sitting is a different task from walking one thesis forward.

A thesis whose screen predates screen review reads an effective accepted screen with no candidate behind it, named on the record rather than papered over with a proposal minted retroactively.
The thesis panel showing one thesis as a four-stage progression — Candidate, Inventory, Screen, Cohort — with the fourth stage marked as the current one and open, the first three collapsed and marked done, and the terminus scope note printed beneath.
One thesis as a staged progression. The current stage is open, the earlier ones are collapsed and reachable, and the sentence beneath states that the journey ends at the cohort by decision — a scope line, not an omission.

The screen is itself reviewable

A thesis does not necessarily have a well-defined screen. So the screen arrives the same way the thesis did — as a candidate you accept, amend or reject:

Discovery — what it reads, and what it refuses

The discovery agent reads this system's own planes and no open web: sixteen declared blocks, of which fifteen are buildable, each declaring its priority, cost, provenance and whether it is untrusted third-party text. The manifest accounts for every block in both directions — included, refused, dropped for budget, not built — and the totals are asserted rather than promised.

The inversion is the heart of it. The economic calendar and the macro containers are deliberately excluded from a per-security assessment, because there is no symbol axis and the per-name elasticity that would give a macro print meaning is not built. Thesis discovery has no symbol axis by design — so the reason those blocks were refused there is the reason they belong here.

Reading a candidate

Conversion records lineage and does not endorse

A converted thesis is born watching and never active — asking for active refuses by name with nothing written. The candidate is not deleted: a conversion record is appended to its chain and every earlier version still says exactly what it said. Authorship is derived per field from the diff, so a thesis you rewrote entirely reports as your assertion and one you changed a word of does not — and where the conversion itself changed a field, it says the software did that rather than crediting you with a change you did not make.

A thesis still cannot graduate to active until it carries a recorded challenge — the strongest opposing case argued from the same data. A token objection is rejected; forcing past the gate is allowed but permanently logged. The risk register lists active convictions nobody has argued against.

A contract change, stated as one. A thesis is now born watching. Until recently create_thesis hard-coded active while the challenge gate lived only on the update path — so every thesis was born active and unchallenged and appeared in the risk register on the same call that created it. Existing stored theses were not touched; the break is on the write path, and both doors carry the sentence saying which contract is in force.
Probabilities are subjective, auditable estimates with a full evidence trail — not validated forecasts. Model-written candidates are the model's judgment, labelled as such at every point, and converting one is an act you take rather than one the system takes for you. Treat all of it as a decision aid.

The Market tab: the heat map & the repricing scan

The 🌐 Market tab is the twin's instrument panel — where the gap between the evidence-supported price and the market's price becomes visible across the whole market at once.

🗺️ Heat map. A treemap of every covered NYSE+NASDAQ name's gap between price and its precision-weighted Model of Models — the residual field. Pick a per-model lens, read it as Δ% or Δ$, scroll back one recorded trading day at a time, and slice it to any cohort. Gap movement is repricing, not dollar flow. Reconstructed history days are labelled as reconstructions.

📡 Repricing scan. A daily scan classifies where models and prices disagree and what to do about each class: names whose gap calls for a strategic-initiative book, cheap tails routed to erosion/regime review, data defects named as defects rather than opportunities, model-maintenance items, and new divergence-ledger activity. Every SI-need row clicks through to that name's Strategic Initiatives editor.

🧬 Valuation structure. The third view: which names are declared operating, hybrid or initiative-principal, and what that declaration does to how their composite is read.

The surface accrues from the daily batch (see Automations); coverage and staleness are disclosed on the map rather than smoothed over.
A sign test is not a floor, and this is what that cost. The composite used to admit any fair value above zero, so a broken denominator of one cent was published as a fair value and the gap was divided by it — one name printed a +20,414% gap, and those values clustered at a few cents regardless of price, which is a broken denominator and not a cheap stock. A value below a stated floor now refuses by name, never clamped and never substituted. The sector means moved as a result: aggregates are taken over the evidenced set with a median published beside the mean and named as the reading to prefer (a gap is bounded at −100 and unbounded above, so its mean is not a location statistic), the incumbent number is preserved verbatim under its own name, and no cell was deleted — every cell carries whether it was evidenced.

Divergence ledger, rerating test & event impact

The twin's discipline is that disagreement between its read and the market's is stored, not discarded:

External text (headlines, transcripts) is treated as untrusted display-only data throughout — it is fenced, never parsed into instructions, and never routes writes. Structured facts drive the loop.

Honest numbers: grades, base rates & implied expectations

Every fitted estimate carries an evidence grade (A–D) computed from real statistics — p-values, confidence intervals, sample size — and capped by method: a regression on a dozen annual points can never grade above B however clean it looks, a sentiment lexicon never above D, chart patterns never above C, while as-filed accounting can reach A. The influence map draws each link with the confidence its evidence earns. Confidence degrades with the evidence automatically, rather than depending on anyone remembering a caveat.

Base rates judge a forecast against what comparable companies actually did: the cohort distribution, where your number sits in it, the share that sustained that level for your horizon, and the empirical fade curve a flat-line model quietly assumes away. Implied expectations run the valuation backwards — what must be true for today's price to be fair — and state the answer against the cohort: "the price requires 21% growth; 4% of comparable companies delivered that."

The top-down half

The bottom-up and top-down prices do not have to agree, and the deliverable is not a second price — it is a dependency list with probabilities, wired to events. The honesty constraint is stated up front: a price constrains a surface, not a point. Price is one number; growth, margin, reinvestment and duration are four, and infinitely many vectors reproduce the same price — so solving each driver separately while holding the others at our assumptions yields four conditional statements about our own model, not "the market's assumptions." What ships today reports the requirement as a price-over-frame ratio with candidate structures and their determinacy; the driver-level inversion inside the intrinsic frames is not yet built, and that gap is stated rather than glossed.

Attention & sizing

The re-underwrite queue scans your holdings and watchlists and pushes the names that deserve attention: structural breaks, new events bearing on a driver the position depends on, worldview drift, a model that has run persistently high or low, or a position never underwritten at all — held on memory alone. An entry is a prompt to re-underwrite, never a signal to trade. Sizing maps the two things the product measures into a weight: the evidence grade caps size, and the expectations read sets the multiplier within the cap. Weights scale down to fit and never up, the do-nothing row is always present, and the mapping is disclosed in full.

Keeping score

Decision records require an expectation, a horizon and a falsifier — the observable that would prove the decision wrong — and the scorecard grades decisions rather than returns, singling out right for the wrong reason as the dangerous column: it pays you to keep a broken process. Rebalance plans state their full cost — trading cost, estimated tax on realized gains — and the hurdle against doing nothing, which is free.

The Community tab: chats & channels

The rightmost tab is Community — a full messaging client built into the terminal, and the transport for sharing research between Open Trading Surface users. It speaks an open standard underneath (XMPP against any server you run, such as Openfire), but the interface never asks you to care: the rail reads Chats, Channels and Accounts, people appear by name with presence dots, channels by their published names, and the plumbing lives in the Accounts panel where plumbing belongs.

The Community tab: the rail with Chats, Channels and Accounts; a channel conversation with a shared chart card; a chart panel with the share control; the Accounts panel
The Community tab. A channel conversation with a shared-chart card (left), the chart it came from with its 📤 share control (right), and the Accounts panel. Conversations are floating panels like everything else in the workspace — a chat sits beside the chart it is about.Captured from a live terminal on 2026-08-18 (5.85.3).

Accounts

Configure accounts under Administration → Chat accounts: server, address, password (stored locally with your other keys, never shown back), whether the connection is encrypted, and connect at start — tick it and the link comes up with the daemon, retries on a declared backoff when the network drops, and catches up whatever it missed, so you never think about "connecting" again. A self-signed server certificate can be allowed per account; the terminal then says, on every read, that the link is encrypted but not authenticated, with the certificate fingerprint shown — a stated trade-off rather than a silent one.

The Accounts panel: connection state with the certificate disclosure, the channel directory with a refresh, People, and Find people
The Accounts panel: the live connection with its encryption disclosure, every channel the server publishes (joined ones marked), join-or-create by address, your contact list, and a search over the server’s user directory.Captured 2026-08-18 (5.85.3).

People

Your contact list renders with live presence; clicking a person opens the 1:1 conversation. Adding someone takes an address or an invite link (a link’s pre-authorization is honoured, so a capable server approves without asking the other side); incoming contact requests appear with Approve/Deny, and approving subscribes back so you can see each other. You can also search the server’s user directory by name, username or email and add people from the results — no address required.

Channels

+ Join or create… in the rail opens the browse dialog: every channel the server publishes, everyone on your contact list, and the user-directory search — pick a channel to join it, pick a person to start a chat. A channel that does not exist yet is created when you enter its address, unlocked for everyone with the server defaults. The directory refreshes as the dialog opens and on demand, so a channel a colleague created a minute ago is there to join.

The browse dialog: channels the server advertises, people, and the directory search
Browse, don’t recall: the channels this server advertises, your people, and a search over the server directory — pick to join or to chat, or create a channel by address on the right.Captured 2026-08-18 (5.85.3).

Conversations

Conversation panels carry the affordances every mainstream client taught: delivery checkmarks on your messages (a ✓ is a confirmation from the recipient’s client, and its absence claims nothing), a “… is typing” line while it is fresh, room subjects and who-is-here, and 📎 Attach when the server runs an upload service — a file goes to the server and its link into the conversation, capped at whatever the server advertises. The log outlives the link: every conversation this installation has ever received stays readable whether or not its account is connected, and on every reconnect the client asks the server’s message archive for exactly what it missed — anchored at the newest message already on disk, deduplicated so a replay is never stored twice. You do not miss messages, and your history does not multiply.

Sharing research, and the hive mind

This is what Surface exists for: sharing live research artifacts between Open Trading Surface users — universally, or between any subgroup. Every chart panel carries a 📤 control. It asks one question — with whom? — a channel or a person, and the chart lands in that conversation as a card: the symbol, the clock, the window, and the chart’s own settings. Whoever clicks Open gets that chart rebuilt as a new panel in their own terminal, from their own data — indicators, overlays, style, window and all.

A shared chart card in a channel: the symbol and clock, an Open button, and the sentence stating reconstruction uses the receiver's own data
A shared chart in a channel. The card is metadata, never data: Open reconstructs the chart from the receiver’s installation, and any series the receiver does not have is refused by name on the chart — never fetched from the sender, never silently dropped.Captured 2026-08-18 (5.85.3).

The availability rule is stated on the card because it is the honest heart of the design: not everyone holds the same data entitlements, so a reconstruction uses your providers and your keys, and a series your installation cannot resolve is refused by name on the chart rather than approximated or silently omitted.

Charts are the first shareable kind, not the last. The share mechanism is a registry — each kind declares how it is collected, how its card reads, and what a receiver’s click does — with theses, screens, cohorts and research notes next: a shared screen arrives as its predicate spec and compiles against the receiver’s own field vocabulary; a shared cohort arrives as its definition and resolves against the receiver’s own universe. A group of researchers on one server becomes a shared surface of charts, screens and theses — each artifact reconstructed locally, each provenance visible.

Why sharing is safe here

Everything that arrives over a chat network is data, never instructions — the conversation panel says so above every scrollback, and the marking travels with each stored message end to end. A sender whose identity the network could not verify is labelled (unverified) beside their name, because a nickname is not an identity. Structured shares cross a closed grammar: a fixed set of reviewed verbs, declared fields only (an undeclared field is refused by name, so nothing can be smuggled inside a share), and a size bound. And the only thing a shared artifact can ever do is be rendered — reconstruction happens when you click, locally, from your own data. For anything beyond rendering, the bar is higher still: an actionable message requires an allowlisted sender and a well-formed envelope and a declared verb and your confirmation, and the only outcome is a proposal you review — nothing in the chat layer executes anything, ever. The allowlist starts empty, which is the safe default: everyone can talk to you; nobody can propose to your machine until you say so.

Portfolios & the transaction ledger

Track what you own and measure it like a professional. Positions are derived from a transaction ledger — nothing is an opaque balance, and the ledger is immutable: a methodology change never edits a transaction, it re-derives what the recorded transactions mean.

The portfolio holdings view listing each open position with shares, average cost, market value, weight and unrealized profit and loss.
Holdings, derived from the ledger. The cost-basis method that produced these numbers is stated on the page.
You sayWhat it does
"Create a portfolio 'Core' and record: bought 100 AAPL at 180, deposited 50k."Sets up the ledger and derives positions, cost basis, cash, weights.
"Import my broker CSV and reconcile any missing dividends."Imports transactions and flags candidates for you to record.
"Show performance vs SPY, and attribute my active return by sector."TWR/MWR, risk stats, and Brinson allocation/selection/interaction.
"Propose target weights that cut risk without dropping expected return much, then a rebalance plan."Runs the optimizer and produces buy/sell deltas (plan only).

The transaction vocabulary

Beyond deposit, withdraw, buy, sell and split, the ledger speaks the whole income and corporate-action vocabulary, each as a deterministic transform inside the one lot engine:

Two laws worth knowing. Income is not price P&L: dividends, interest and withholding aggregate separately and enter position P&L as income, never blended into price appreciation. And a recorded income event must be possible: a dividend on a symbol the ledger never held on or before that date refuses at record time.

Delisting is the instructive one: it transforms nothing — your lots and basis are untouched — but the position becomes unpriceable, and every valuation of the book refuses naming the delisting date until you resolve it with a sale, a tender, or a worthless record. That is refusal-over-default applied to your own money.

Detection assists, never corrects. Your ledger is compared against provider split history in both directions, and missing dividends are detected — but every finding is flagged for investigation, never applied. A flag is a question, not an edit.

Sorting & the house table

The portfolio surface is nineteen tables across eight views, and until recently not one of them sorted. They do now — and the point is that it was not done table by table.

A table declares its columns, and the sort arrives with the declaration

A sort added per table is the curated list this house says rots: it is exactly how three tables end up sorting and two silently do not. So a table declares its columns through one door, the sort comes with the declaration, and the build fails on any portfolio renderer that emits a heading outside it.

Two tables refuse to sort, by name

The cash-flow reconciliation is an ordered identity, not a list:

opening + contributions − withdrawals + income − fees ± market change = closing

Reordering it leaves it numerically correct and meaningless, so it refuses and offers no affordance at all. Three footer blocks — the decomposition's residual, the contribution-to-risk total, the realized table's short/long subtotals — were ordinary rows that would have sorted into the middle of their own bodies, and are now footers. A position's tax lots are that position's detail and travel with their parent row rather than being sorted away from it.

One house table

Every table in the terminal now has a maximum width of 1,120px, drag-resizable columns, and sortable headings where the surface has been migrated. The width is measured, and what it was measured on is the decision: it is the natural width of the widest table of bounded content — the twelve-column holdings table, 990px at a 31-character name and 1,117 at 50. A free-text column is bounded by wrapping, not by the table, so deriving a ceiling from a sentence would let one long title decide the width of every table in the terminal. The calendar grid refuses the ceiling, in the stylesheet next to the rule it refuses.

A resize pins every column and the table takes the sum of its columns — before that line existed, a −120px drag moved the column 95px.

What is not yet migrated is declared rather than implied. The max width applies to every table on day one, because it is a stylesheet rule that needs no adoption. Sorting and adjustable widths arrive with a column declaration, and a survey found 584 table headings across 53 functions outside that door. That population is a dated debt policed by a ratchet — a count that grew fails the build, and a count that fell fails too, because paying it down has to be a visible edit rather than a number that drifts.
A sort is the panel's own state: it survives a re-render and a reload, and it deliberately does not travel in an exported layout — stated on the payload — because the portfolio surface is a pinned surface rather than a workspace panel. sort_portfolio_table runs the same operation the click runs.

Cost basis & tax lots

Five methods, from one lot engine — there is no second derivation path anywhere in the system:

MethodWhat it consumes
FIFO (default)Oldest lot first.
LIFONewest lot first.
HIFOHighest unit cost first, ties broken oldest-first.
Average costEvery sold share prices at the pooled average basis; shares and holding periods are consumed FIFO. Open lots then display the pooled average, so basis is conserved.
Specific identificationExactly the lots the sale names, and nothing else. The engine never picks lots for you.

Every realized event records its shares, proceeds, basis, P&L, open and close dates, days held, term, the method used, and the buy and sell transaction ids. Term is long when shares were held more than a year: a sale 365 days after the buy is short, 366 is long. That boundary is deliberately not tunable.

Performance does not move when you change method

Performance is flow-based, so TWR and MWR are method-invariant: the daily equity series reads transactions and prices, and basis never enters it. Changing the method moves the composition of realized versus unrealized — and the short/long-term tax split — but their sum, your cash, total equity and the equity curve do not move by one byte. That is why a method change is safe to make: the published performance record cannot shift underneath a methodology choice.

Set it with set_cost_basis_method or the terminal's method selector. The change is journaled and returns the before-and-after realized short- and long-term totals, so you read the effect before you believe it. A sale may also carry a per-sale override.

What it refuses

Validation runs twice: when you record the transaction (the candidate ledger is derived whole before the transaction is accepted) and again at derive time, because an imported ledger can carry a bad specific-ID.

Not built, and named so its absence reads as a decision: wash-sale adjustment, tax-rate application, jurisdiction rules and after-tax return series. The layer states realized short- and long-term dollars; it computes no tax owed. Also out of scope: multi-currency ledgers, options, bonds and derivatives, and any GIPS compliance claim.

Performance, attribution & risk

The portfolio performance view showing an equity curve against its benchmark, a period returns table and a panel of risk statistics.
The equity curve against a total-return benchmark, the period table, and the risk statistics.

The benchmark, and a dated methodology break

The benchmark is total-return: declared dividends reinvested on their ex-dates. Before, it used the tracking ETF's price series — disclosed, but a price series excludes roughly 1.3–1.5% a year of dividends, so every excess-return reading was systematically flattered by about that much. Honest disclosure of a flattering construction is still a flattering construction.

Replacing it moved the published excess-return series downward, so every payload and page that defines the benchmark carries the sentence "TR benchmark since 4.84.0; prior reports used a price proxy." Where dividend history is unavailable the builder falls back to the price series and the disclosure says price-only, with the bias named. Where the benchmark is proxied by a tracking ETF rather than reconstructed from full membership, the report names the proxy — refusing to disclose a proxy is the defect; the proxy itself is honest practice.

No other published number moved. Net TWR is the same equity curve it always was; gross TWR, income, the period table and the reconciliation are new disclosures, not adjustments.

The reconciliation that must sum

Every report window prints a client-statement reconciliation:

opening value + contributions − withdrawals + income − fees ± market change = closing value

The closing value is derived from the daily equity series and the right-hand side is derived independently from the raw ledger and prices — two routes to the same number. A residue beyond one cent throws, naming the residue. So a reconciliation table you can see is, by construction, a balanced one; if it cannot balance, the report refuses rather than printing. Return of capital gets its own line, because it is neither income nor a contribution. Trade commissions live inside market change; the fees line is explicit fee transactions — and the table says so.

The portfolio allocation view showing composition by sector and by position with weights and concentration measures.
Allocation and concentration — the inputs the ex-ante risk panel decomposes.

The PDF report

The Report view assembles a full institutional tear-sheet and offers Export PDF. The subject can be a portfolio (full treatment, with flows) or a cohort treated as a synthetic buy-and-hold book — and a synthetic book never presents a money-weighted return, because money-weighting a book with no money is not a number. The cover discloses which kind it read.

Contents, in order:

  1. Cover — the lead number (window or since-inception TWR versus benchmark) dominating the page, a composition donut with total equity in the center, the KPI strip, the subject kind, the cost-basis method, and the benchmark with any proxy named.
  2. Performance — equity curve versus benchmark, the period table, the risk statistics, the gross/net-of-fee line, and the cash-flow reconciliation for the window.
  3. Attribution — Brinson bars by sector, the allocation/selection/interaction table, linked monthly effects, and the factor panel beside them.
  4. Risk & stress — ex-ante forecast volatility and tracking error, contribution to risk, concentration, and the stress table.
  5. Holdings — one row per open position with weight, return, contribution and a valuation percentile; this is the drill-down index. Fully closed names appear in a compact Realized section with method, term and short/long subtotals — a closed position is not a holding, so it never ranks the drill-down order.
  6. Per-position pages, in contribution order — identity strip, the vertical (valuation percentile, own-history band p10–p90, divergence entries), analyst context and catalysts, your thesis note quoted with its date or the honest sentence that none is on file, pros and cons each naming their source, and a descriptive positioning reading with the reasoning shown.

Two routes, one assembly. From the terminal, Export PDF streams the bytes over HTTP as an ordinary browser download and writes nothing to disk; a refusal answers as plain text so the sentence reaches you verbatim rather than being saved as a broken PDF. From the agent, export_portfolio_report writes the file into your state directory and returns the path plus a page inventory — and never auto-opens it. get_portfolio_report returns the same data model as JSON.

The PDF writer is zero-dependency: PDF 1.4, built-in fonts, vector graphics only, page tree and cross-reference table assembled by hand. No raster, no external fonts, no library. It formats; it never computes. A pie or donut is admissible only for a part-of-whole with at most five non-aggregate slices and no negative weight — otherwise it falls back to horizontal bars, because a part-of-whole picture cannot carry a sign. Comparisons use paired bars on one shared zero axis, never dual axes.
What the report refuses, and why. Attribution refuses below a constituent-mapping coverage floor. Ex-ante risk refuses below a minimum observation count, naming the count — a young ledger refuses honestly; nothing is wrong except youth. The factor panel refuses when fitted loadings cover too little of the book, with the coverage number. Economy-scenario stress rows always refuse, because the chain from a saved economy scenario to a per-position spot price does not exist in this system and asserting a pass-through would put an opinion where a measurement belongs. Drill-downs are capped at 40 positions, with the overflow named in the holdings table as "not drilled" rather than silently dropped. And no output is advice: hold/add/trim language is a description of what the vertical says, with the reasoning shown, under the standing disclaimer.

Where a holding came from

A position should trace back to a view. That only means something if the chain is recorded rather than reconstructed, so building a model portfolio from a thesis snapshots the whole lineage: the thesis it came from, the compiled screen exactly as it ran, the run id, and every pick with the rank score that earned it a place and the weight it received.

Open a portfolio and you'll see the breadcrumbThesis › Screen › Book — each hop clickable, and the thesis links forward to the book it produced. Ask "why do I own these names?" and the same chain comes back.

Links are by id, not name, so renaming never breaks the chain. The screen spec is a snapshot, because a compiled screen is a function of the thesis's current tilts — re-deriving it after you edit the thesis would answer "what would this screen for today", not "why do I own this".

The view above the thesis

When you create a thesis you cite the worldview domains it stands on — growth, inflation & rates, geopolitics, industrial policy, technology, demographics, energy, credit — and Open Trading Surface snapshots what each of those said at the time. That snapshot is what makes drift visible: if you formed a thesis when inflation was deteriorating and it has since turned improving, the breadcrumb flags it. Stale domains are flagged separately, because not having looked is different from having seen a change.

Books built before this existed can be reconstructed by name-matching, but they are labelled inferred and carry no screen spec — none was ever stored, and inventing one would misrepresent an audit trail. A portfolio you created by hand honestly reports no lineage rather than guessing at one.
The eight sections that follow are the walking tour. The full event-processing guide is the reference: the architecture diagram, the complete event vocabulary with which kinds are live, which are written only on demand and which 21 are dormant and can never fire, the condition grammar, every backtest refusal, and a worked end-to-end strategy.

Reach: what can honestly be backtested

Administration → Reach answers one question before you ask any other: what could this installation honestly backtest, and over what span? It ships deliberately before the engine that uses it, so the limits are visible before expectations form.

The Reach panel listing predicate families with the history that exists for each, how far back a backtest could honestly run, and refusal sentences naming the missing point-in-time record.
Reach. Where the record does not exist, the answer is a refusal naming what is missing and what would fix it — not a number.

Every number on it is a live read of your own stores. It is the panel to consult before writing a condition — which is why the two sit one click apart, the disclosure ordering made physical.

Why a record cannot be conjured after the fact. Several things this system needs can only accrue forward: the dated cohort-membership ledger (whether a name was in a cohort on a given date), the monthly universe snapshot, and the point-in-time panel archive. What is not accruing today cannot be backtested tomorrow — so the honest thing is to say so, start the clock, and refuse windows that fall before it.

Typed events

Everything the engine reasons over is a typed event in one vocabulary, published as plane.family.name and built from the producers' own exported lists rather than a hand-maintained catalog. Recording an unregistered kind throws.

The two timestamps

Every event carries when it happened and when it was first knowable to us. A 10-K filed in February covering the year ended December happened on 31 December and became knowable in February. A backtest at a given bar may only see events that were knowable by that bar's close; a chart or a report that wants the economic date reads the other one. Where a producer genuinely cannot establish knowability it writes none and marks the event unreconstructable, and the backtester treats that as never knowable — it refuses rather than dating a quarter by its period end.

The Events panel showing a paged feed of typed events, each row carrying its event kind, both timestamps with the gap between them labelled, and its identifying predicates.
The event feed. Both timestamps on every row, with the divergence labelled rather than signed.

What writes events

The feed pages from the newest day file backwards rather than sweeping a range, because the log has no retention policy by design and a surface built on a range read would get slower every day the engine ran. The counts describe the page and say so. A filter matching nothing stops at a bounded number of files and returns its cursor with exhausted: false — because "no more matches" and "I stopped looking" are different answers.

The condition language

Administration → Conditions is where you compose what you want watched. A condition is canonical JSON with a text sugar that round-trips, composed through the same one selector over the same taxonomy the screener uses. It joins the cross-sectional predicate language to the temporal series language with real CEP operators:

Durations carry their clock. 20d is twenty trading days; 20c is twenty calendar days. A bare number is a parse error that names both — the ambiguity is not resolved for you.

Each condition is checked against the reach report and shows its point-in-time verdict beside it. A leg with no point-in-time value refuses by name, naming the leg, the rule and the plane.

Rules & the tick

Administration → Rules. A rule is a named, versioned condition with a subject scope resolved live. Roughly: a condition asks a question; a rule asks it of particular names, on a schedule, and remembers.

Your existing classic alerts can be migrated to rules. Seven of the nine legacy types moved; two were refused by name rather than approximated — one because the event kind it needs does not exist, one because the kind exists and nothing writes it, so the rule could never fire — and both were left running on the old path rather than silently dropped.

Actions & the paper book

A firing can cause something. Nine typed verbs are published as one registry that both the tool surface and the terminal picker derive from:

VerbWhat it does
propose_buy · propose_sell · propose_sizeWrite a proposed transaction into a paper book.
flagRecords a flag with a severity and surfaces it in the feed. No book effect.
annotateWrites a research note through the one note store.
notifyDelivers the firing to the agent inbox — the explicit form of what a noisy rule does implicitly, so a rule can be silent by default and loud on one branch.
run_playbookQueues a stored playbook.
emitWrites a further typed event (with a cascade depth cap that refuses by name).
set_stateSets a named per-subject state.

Actions attach to the rule, not to its condition. That is not a detail: the condition's fingerprint is the per-subject state key, so folding an action into it would mean that retagging a proposal re-arms the rule on every name it watches. An action edit therefore inherits state.

The paper book — proposals only

The only consequence that touches a book is a proposed transaction in a paper book: a real ledger, in a separate file, appended through the same builder your real trades go through, so the one lot engine validates a proposal exactly as it validates a real trade. It is measured by the existing report machinery — no performance code was written for it — and gated by the same reconciliation. There is one book per strategy, because comparisons need independence.

There is no execute verb, no order verb, and no promote path anywhere in this system. That is not a policy that could be relaxed; it is enforced structurally in four places: a different file with an executable guard; a write path that resolves targets only inside the paper store and refuses a real portfolio by name (a bare "not found" would read like a typo and invite a retry); only the pure functions of the portfolio engine reachable, with a test that fails the build on any store-writing verb, any execute/order/broker verb, and any promote path; and paper: true stamped at the write and marked on every payload. The PDF cover prints PAPER BOOK — PROPOSALS ONLY.

A size that cannot be computed skips and records the binding constraint rather than defaulting. A propose_sell on a name the book does not hold refuses by name — there are no short lots in this engine.

Strategies

Administration → Strategies. A strategy is where firings stop being independent and start competing — for cash, for slots, for the sector budget, and for the liquidity the tape will bear. A stated policy settles it instead of iteration order.

A strategy binds a universe, a rule set with roles (entry · exit · adjust · veto — the role belongs to the strategy, so one rule can be one strategy's entry and another's veto), a policy, and one paper book.

Sizing

Seven methods, each stating its budget basis: fixed_cash, fixed_pct, equal_weight, grade_ceiling, risk_parity (inverse-vol weights scaled against the book's target volatility and capped at fully invested, because this engine does not lever), atr_risk (a fixed fraction of equity put at risk across concurrent positions) and from_targets (the strategy's stored targets through the rebalance planner, read and never re-derived).

A stop distance used by atr_risk is a sizing input, not an order. This engine places nothing and there is no stop-loss verb; the ATR distance is the denominator of the risk budget and nothing else. A rule that wants to exit on that level says so as an exit condition.

Constraints

Constraints compose, and the binding one names itself, with every other that would also have bound listed beside it. Ceilings shrink and say so; floors and counts refuse. A minimum position size is a refusal, not a rounding. A name with no sector is refused when a sector cap is set. An average-daily-volume cap refuses the order rather than inventing a partial fill.

Ranking, cash, and the daily loop

A rank field is required — if the strategy does not say how to rank its candidates, the answer is hash iteration order. Nulls rank last, and exact ties are broken by a declared rule that says so on the row. Cash yields the three-month bill accrued actual/360 where a FRED key is configured, and zero where it is not, with the source disclosed either way.

The loop walks each day in this order:

  1. Corporate actions
  2. Cash accrual, before any trade
  3. The knowability cut
  4. Vetoes
  5. Exits before entries — an exit frees cash and a slot
  6. Rank
  7. Size and constrain
  8. Fill at the bar close, with costs booked per side as a fee
  9. Rebalance
  10. Mark, and reconcile — which throws above a one-cent residue and has no flag to skip it

The loop is driven by an injected clock and reads the wall clock nowhere inside its walk. That is the single most important structural decision in the layer, because the backtester replays this same loop rather than reimplementing it — and two implementations would make every backtest number a claim about code the live book never ran.

Editing a strategy is a methodology break, not a reset. An edit that changes how positions are formed (universe, rules, policy) bumps the strategy version and writes a dated methodology break — and does not reset the book. A rule's per-subject state is a claim about a condition and means nothing under new semantics; a book is a record of proposals actually made at prices that actually printed, and resetting it would delete the only out-of-sample evidence the strategy has. A performance window spanning a break says so, because an equity curve across a policy change is two measurements.

The backtester & deflation

The standing posture is stated before anything runs: bias toward refusing over producing a number. A backtest that declines to run is a result; a backtest that quietly used tomorrow's data is a lie that will be believed.

A backtest is not a simulator. It is the same daily loop the live book runs, driven by a replay clock, writing real ledger records through the same builder into a paper book of its own — never the strategy's live forward record, which is the out-of-sample evidence hindsight cannot manufacture — reconciled on every replayed day, and measured by the existing report engine. The backtester computes no return, no Sharpe, no drawdown and no attribution of its own.

What it refuses, and why

The one cross-sectional hazard the tape can answer is answered: a name whose first bar postdates a given day is excluded from that day's cross-section and counted.

Multiple testing, made unavoidable

Every run registers its trial before it reports. A run that registered only on success would count the winners and forget the losers. Every result carries its trial count on the face of it, and the Deflated Sharpe Ratio is computed per observation; below a minimum trial count the engine reports a Šidák correction instead and says which it used rather than fabricating one.

Runs can be pruned (dry-run by default). The trial record cannot — there is no prune verb, no cap and no retention window on it anywhere, and the asymmetry is the point.

"How much of it was one name?"

A backtest whose result comes from one name is not a strategy; it is a story about that name. So the drop-top-names re-run replays the same strategy over the same window with the top contributor removed from the universe, and reports whether the result survives. It is the honest complement to the deflation: that one asks how many things you tried, this one asks how much of it was one lucky name.

What this actually found. On the reference test strategy over 2019–2024, the reported result was +53.36% with a Sharpe of 0.54 and a maximum drawdown of −26.55%, reconciled to zero residue on all 1,510 days. Dropping the single top contributor left +18.74% — about a third of it retained. Dropping the top two left −7.70%: the three remaining names lost money over six years while the headline said a 53% gain. That is the entire argument for this feature.

Events, Firings, Runs & Backtests

Four panels make the engine visible. They add no engine capability — every payload behind them already existed and was already tested. They are the surfaces that were missing.

Events

A browsable stream of the typed event log, described under Typed events. Every row carries both timestamps with the divergence shown as a gap, the reconstructability verdict, and identifying predicates that are the registry's own declared payload keys. The filter is the one selector over the registry's vocabulary — publish a new event kind and it appears in the picker with no client edit.

Firings

The firings and the non-firings, which is what the accrual record was designed for. It keeps three states apart:

Coverage counts days at every level including per subject — 40 names over 3 days is 120 rows and three days. The live per-subject state shows open partials and live timers, so a sequence rule mid-flight is legible, and each firing joins to the typed event carrying its two timestamps.

Runs

The strategy loop's most recent day, laid out in the loop's own order. A day where nothing happened prints "RAN, NOTHING FIRED" — a recorded success, and a different answer from silence. A strategy with no stored day renders NO RECORD. An empty panel and a day walked in full are opposite facts that would otherwise look identical.

Backtests

The results surface. It computes nothing — every figure is read off the run store, the paper book's own ledger, the existing report and the search ledger. Three properties the renderer cannot drop:

A detail page runs markers → honesty block → counterfactual → numbers: the deflation with which statistic was used, the concentration, every refusal with its leg, the caller's own selection bias, the generations spanned, the reconciliation residue, the holdout; then the drop table with both rankings and the sentence saying why they can disagree, the cost stated before it runs and the floor and cap refusing before anything replays; then the equity curve against the benchmark (an inline vector chart, no library — and where the report refused, the loop's own reconciled marks with no benchmark line said rather than drawn), the period table, the blotter, and the cash-yield source.

Alerts, subscriptions & the calendar

The 🔔 Alerts tab has three sub-sections, all fed by the same typed-alert machinery:

Feed — the attention-focusing stream. Typed notifications (divergence activity, rerating candidates, SI gates approaching, band crossings, scan findings, executive changes, price-rule firings) grouped by day then type, with per-type count chips, aggregated day summaries, the deciding subscription layer on hover, and click-through to each type's evidence surface. Classic price rules live in a panel inside the feed, with playbooks.

Subscriptions — deliberate control over which alert types fire on which securities, cohort-first: pick a cohort, toggle types on or off as standing rules new members inherit, set defaults, and drill into per-symbol overrides.

Calendar — one time axis for everything dated: typed alerts, SI gate expected resolutions (overdue flagged), corporate events, ingest intake items, divergence-ledger entries, roster changes, economic releases and your own calendar events — as a list or a visual month grid with day drill-down, type filters, cohort/symbol slices, and click-through per row.

This calendar and the Macro-economic → Calendar panel are two views of the same records, not two feeds: the economic rows are read from the same store under the same ids, so the same event appears on both surfaces and clicking through resolves to the same thing.
The classic alert machinery and the rule engine coexist deliberately. Seven of the nine legacy alert types have a rule equivalent you can migrate to; the other two were refused by name rather than approximated, and continue to run on the old path.

Executive profiles

Every Key Executives row in Company Information opens a full profile page: the provider record (title, pay, tenure), roster-change history as recorded (first seen, title changes, departures — with the honest caveat that earlier tenure is unobserved, not absent), compensation history, this person's insider-trade rows (deterministic name matching, rule stated on the page), career across companies cross-referenced from recorded rosters and Form-4 filings, a Wikipedia summary fetched on demand (exact-name match only — ambiguity is stored as ambiguity, never guessed), candidate news mentions (fenced, with the common-name caveat), and an Analyst report section rendering agent- or owner-authored notes. Roster changes are detected daily and flow into the scan report, the calendar, and subscription-gated alerts.

The marketplace & images

The 🛍 Marketplace tab browses and installs images — shareable packs of the twin's state: reference data (universe, recorded surface days), authored model books, and configuration. Installs are verified against the catalog checksum and recorded in a provenance ledger; imported records carry their origin and never masquerade as your own history. Versioned artifacts are immutable — a published version never changes; new versions append. The free OTS Starter Image bootstraps a fresh install with the full NYSE+NASDAQ surface, strategic-initiative books, and workspace configuration — no keys, no portfolios, no personal research. You can also export your own packs to share.

Research notes & memory

Open Trading Surface has a shared memory: research notes tied to symbols, portfolios, or theses that persist between sessions, so you and the agent build on prior work instead of starting over. Ask the agent to "write your take on this name to research notes," or "what did we conclude about MSFT last month?" Notes are also read by the PDF report's per-position pages and can be written by a rule's annotate action.

The News panel for a security showing a dated list of headlines with publisher and link, rendered as untrusted display-only text.
News as it reaches you — and, underneath, as typed events carrying only predicates that can be checked.

Giving feedback

Open Trading Surface is being shaped by the people who use it. The easiest way to report a bug, request a feature, or share an idea is to just tell the agent — "report a bug: the rebalance view didn't load," or "I have an idea: add options analytics." It composes a ready-to-post GitHub report prefilled with your details and version for you to submit (it never posts on your behalf). You can also browse and follow every issue on the Support page, join Discussions, and optionally enable a local-only usage journal ("turn on usage reporting") to attach to a report.

Updating

Update checks are non-blocking and only request a version number — they never transmit your data. Your settings, portfolios, theses, notes, panels and workspace arrangement persist across updates. After an upgrade, if a browser tab is still running JavaScript from the older build, the terminal discloses it rather than behaving strangely.

Uninstalling

Your data & privacy

Everything lives in ~/.ots/ on your machine: keys, portfolios, watchlists, theses, notes, company models, alerts, rules, strategies, paper books, the event log, the trial record, an audit journal, and test reports. Nothing is uploaded to any server operated by the publisher. The only outbound calls are read-only requests to your chosen data providers (FMP, FRED, SEC EDGAR), a version-number check to the update channel, and — only if you configure a key and run a Perspective — a request to the AI vendor you chose. Your workspace arrangement lives in your browser's local storage.

Troubleshooting