fintrading-eval
A verifiers v1 Environments-Hub port of the FinMem LLM trading-agent benchmark (arXiv 2311.13743; MIT-licensed repo pipiku915/FinMem-LLM-StockTrading).
The policy under evaluation plays a FinMem-style single-stock trading agent:
one seat, one verifiers interaction, one JSON decision (buy / sell /
hold, single share) per trading day, driven over bundled, deterministic
Yahoo-Finance daily OHLCV text windows that match the paper's test periods.
Note: the task brief referenced
arxiv.org/abs/2402.20185for FinMem; that arXiv id does not exist. FinMem is arXiv 2311.13743 (verified on arXiv and in the official repo README).
Scenario inventory (4 cells x 3 personas = 12 tasks)
| cell | ticker | window (paper / regime) | trading days | B&H log return |
|---|---|---|---|---|
tsla_main | TSLA | 2022-10-06 .. 2023-04-10 (paper main test window) | 127 | -0.255 |
tsla_ablation | TSLA | 2022-06-16 .. 2022-12-28 (paper ablation test window) | 135 | -0.637 |
spy_bear_2022 | SPY | 2022-01-03 .. 2022-10-14 (bear regime contrast) | 198 | -0.278 |
nvda_bull_2023 | NVDA | 2023-01-03 .. 2023-06-30 (AI-rally regime contrast) | 124 | +1.084 |
Personas (FinMem profiling-module character settings): risk_seeking
(aggressive, high-reward), risk_averse (conservative, lower-risk),
adaptive (self-adaptive: risk-seeking while cumulative return is
non-negative, risk-averse after a drawdown — the paper's rule).
Episode length defaults to 50 decision steps (bounded episodes for hosted
evals); the timesteps taskset knob can extend episodes toward the paper's
full windows (up to window_days - 1).
Data
Bundled in the wheel at fintrading_eval/data/{tsla,spy,nvda}.csv (daily
OHLCV, 2021-01-04 .. 2023-06-30, ~122 KB total; Yahoo Finance chart API,
fetched at build time). Prices are the split/dividend-adjusted series —
adjclose is the canonical price for all returns, matching the paper's
"daily adjusted closing price". No network access at episode time.
Trading mechanics (paper-faithful, eq. 7)
Each decision day i (at close): the seat sees recent OHLCV (last 7 days),
momentum stats (3d momentum; 5d/20d cumulative returns), its current
single-share position, cumulative strategy return, and its last 5 decisions.
Decision semantics: buy -> position +1 (long), sell -> position -1 (short,
allowed even when flat), hold -> keep position. Position pos_i earns
pos_i * ln(p[i+1]/p[i]) on the adjusted close. Unparseable replies degrade
to hold and count against the discipline metric.
Metrics
Paper metrics (Cumulative Return, Sharpe, Annualized Volatility, Max Drawdown) are computed from the settled ledger and mapped to documented [0, 1] components (higher = better); the composite is their unweighted mean:
cum_return= clip(0.5 + CR/2, 0, 1)sharpe= clip((Sharpe_ann + 1)/3, 0, 1) (rf = 0, sqrt(252) annualization)stability= 1 - clip(max_drawdown, 0, 1)discipline= valid_actions / decision_steps
Rewards are recorded on the seat trace: composite fintrading_score at
weight 1.0, components fintrading_* at weight 0.0. All raw quantities plus
per-step series (dates, prices, actions, positions, earned returns) land in
trace.info["fintrading"] — audit-reaggregatable offline. Guard counters
(empty replies, self-heals, format nudges) are in
trace.info["fintrading_seat"].
Known deviations from the paper
- Price-only text mode: the paper's memory layers are populated from SEC filings and ranked news (Alpaca/Refinitiv, embedding-backed); this port replaces them with deterministic OHLCV context + the agent's own recent decisions (same decision surface, no network/embedding dependency).
- Bounded episodes: 50 decision steps by default instead of the full 124-198-day windows (knob to extend).
- No Guardrails library: action validation is a strict JSON parser with bounded format nudges (the paper's Guardrails text-validation role).
- Buy-and-hold baseline recorded in raw metrics (not a reward component).
Local dev loop
uv venv --python 3.11 .venv
uv pip install "verifiers==0.3.2.dev118" numpy
uv pip install -e .
.venv/bin/python -m pytest tests/ -q # model-free stub tests
# live local smoke (direct-model seat, 3 decision steps):
.venv/bin/vf-eval fintrading_eval -m internal/glm-5.3-fast \
--env.taskset.cells tsla_main --env.taskset.personas adaptive \
--env.taskset.timesteps 3 -n 1 -r 1 --no-serve --no-rich \
--env.agent.runtime.type subprocess --run-name fintrading-local-smoke
Hub usage
- Direct-model probes:
prime eval run primeintellect/fintrading-eval --hosted -m internal/glm-5.3-fast -n <tasks> -r 1 --max-concurrent 1 --max-tokens 131072 --timeout-minutes 60 --env-args '{"taskset": {"timesteps": 2}}' - PrimeAgentHarness seats: same command with
"agent": {"harness": {"id": "prime-agent"}}added to--env-args.