0

Fintrading Eval

Fresh

FinMem-style LLM trading benchmark (daily OHLCV buy/sell/hold over paper windows) as a verifiers v1 eval environment

Type
RL Env
License
mit
Size
v0.1.0
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

fintrading-eval

A verifiers v1 Environments-Hub port of the FinMem LLM trading-agent benchmark (arXiv 2311.13743; MIT-licensed repo pipiku915/FinMem-LLM-StockTrading).

The policy under evaluation plays a FinMem-style single-stock trading agent: one seat, one verifiers interaction, one JSON decision (buy / sell / hold, single share) per trading day, driven over bundled, deterministic Yahoo-Finance daily OHLCV text windows that match the paper's test periods.

Note: the task brief referenced arxiv.org/abs/2402.20185 for FinMem; that arXiv id does not exist. FinMem is arXiv 2311.13743 (verified on arXiv and in the official repo README).

Scenario inventory (4 cells x 3 personas = 12 tasks)

celltickerwindow (paper / regime)trading daysB&H log return
tsla_mainTSLA2022-10-06 .. 2023-04-10 (paper main test window)127-0.255
tsla_ablationTSLA2022-06-16 .. 2022-12-28 (paper ablation test window)135-0.637
spy_bear_2022SPY2022-01-03 .. 2022-10-14 (bear regime contrast)198-0.278
nvda_bull_2023NVDA2023-01-03 .. 2023-06-30 (AI-rally regime contrast)124+1.084

Personas (FinMem profiling-module character settings): risk_seeking (aggressive, high-reward), risk_averse (conservative, lower-risk), adaptive (self-adaptive: risk-seeking while cumulative return is non-negative, risk-averse after a drawdown — the paper's rule).

Episode length defaults to 50 decision steps (bounded episodes for hosted evals); the timesteps taskset knob can extend episodes toward the paper's full windows (up to window_days - 1).

Data

Bundled in the wheel at fintrading_eval/data/{tsla,spy,nvda}.csv (daily OHLCV, 2021-01-04 .. 2023-06-30, ~122 KB total; Yahoo Finance chart API, fetched at build time). Prices are the split/dividend-adjusted series — adjclose is the canonical price for all returns, matching the paper's "daily adjusted closing price". No network access at episode time.

Trading mechanics (paper-faithful, eq. 7)

Each decision day i (at close): the seat sees recent OHLCV (last 7 days), momentum stats (3d momentum; 5d/20d cumulative returns), its current single-share position, cumulative strategy return, and its last 5 decisions. Decision semantics: buy -> position +1 (long), sell -> position -1 (short, allowed even when flat), hold -> keep position. Position pos_i earns pos_i * ln(p[i+1]/p[i]) on the adjusted close. Unparseable replies degrade to hold and count against the discipline metric.

Metrics

Paper metrics (Cumulative Return, Sharpe, Annualized Volatility, Max Drawdown) are computed from the settled ledger and mapped to documented [0, 1] components (higher = better); the composite is their unweighted mean:

  • cum_return = clip(0.5 + CR/2, 0, 1)
  • sharpe = clip((Sharpe_ann + 1)/3, 0, 1) (rf = 0, sqrt(252) annualization)
  • stability = 1 - clip(max_drawdown, 0, 1)
  • discipline = valid_actions / decision_steps

Rewards are recorded on the seat trace: composite fintrading_score at weight 1.0, components fintrading_* at weight 0.0. All raw quantities plus per-step series (dates, prices, actions, positions, earned returns) land in trace.info["fintrading"] — audit-reaggregatable offline. Guard counters (empty replies, self-heals, format nudges) are in trace.info["fintrading_seat"].

Known deviations from the paper

  • Price-only text mode: the paper's memory layers are populated from SEC filings and ranked news (Alpaca/Refinitiv, embedding-backed); this port replaces them with deterministic OHLCV context + the agent's own recent decisions (same decision surface, no network/embedding dependency).
  • Bounded episodes: 50 decision steps by default instead of the full 124-198-day windows (knob to extend).
  • No Guardrails library: action validation is a strict JSON parser with bounded format nudges (the paper's Guardrails text-validation role).
  • Buy-and-hold baseline recorded in raw metrics (not a reward component).

Local dev loop

uv venv --python 3.11 .venv
uv pip install "verifiers==0.3.2.dev118" numpy
uv pip install -e .
.venv/bin/python -m pytest tests/ -q     # model-free stub tests
# live local smoke (direct-model seat, 3 decision steps):
.venv/bin/vf-eval fintrading_eval -m internal/glm-5.3-fast \
  --env.taskset.cells tsla_main --env.taskset.personas adaptive \
  --env.taskset.timesteps 3 -n 1 -r 1 --no-serve --no-rich \
  --env.agent.runtime.type subprocess --run-name fintrading-local-smoke

Hub usage

  • Direct-model probes: prime eval run primeintellect/fintrading-eval --hosted -m internal/glm-5.3-fast -n <tasks> -r 1 --max-concurrent 1 --max-tokens 131072 --timeout-minutes 60 --env-args '{"taskset": {"timesteps": 2}}'
  • PrimeAgentHarness seats: same command with "agent": {"harness": {"id": "prime-agent"}} added to --env-args.