‖ Jev Trader
← Back to FAQ

Understanding Jev Confidence Before Evaluating a Trading Strategy

A concentrated model answer can look persuasive. That still leaves several research questions: Does the probability match observed outcomes? Can the simulated order fill? Does the result survive costs and a later test period? These questions need separate evidence. This guide outlines an evaluation method for developers and researchers. It reports no backtest results and makes no trade recommendations.

1. Separate the quantities

For Choice, probabilities assigns values across defined options, summing to one; choice identifies the largest. The options and their descriptions define what the answer means. See the official Choice documentation.

TypeSafe's confidence is a 0 to 1 summary derived from the distribution's shape. A concentrated distribution indicates greater certainty. It is not interchangeable with an option probability or a measured success rate. The reviewed page does not specify the exact formula. Choice and Score carry this field; Noul does not. See TypeSafe's confidence documentation and HTTP answer reference.

An observed win rate is different again. It counts profitable completed trades under a stated accounting rule. Classification accuracy counts correct labels. A directional forecast can be correct while an order remains unfilled or loses after costs. The existing Jev Score guide explains output interpretation; here the focus is testing against outcomes.

2. Define an outcome before collecting predictions

Write a label that another researcher can reproduce. Specify the market, price source, observation time, horizon, treatment of unchanged prices and missing data. Keep future information out of the input snapshot.

Illustration only, not a model response or measurement: a research question asks whether the mid-price will be strictly higher after 30 seconds. Its two outcomes are higher and not_higher, with illustrative probabilities 0.70 and 0.30. This does not establish a 70% profitable-trade rate. No confidence value is calculated here.

Keep forecast labels separate from action labels. The current trading prompt combines future direction with post-only execution and allowed-side constraints. Its buy weight cannot automatically be treated as a calibrated probability of a future price rise. Freeze a prediction-only question for that experiment, and evaluate execution separately. Record the question, criteria, actual model version, timestamps and raw answer before observing the label.

3. Test calibration rather than assuming it

On held-out observations, group forecasts for the same event into probability ranges. For each range, compare the mean predicted probability with the observed event frequency. Show counts alongside the comparison. This is a reliability diagram, as described in scikit-learn's calibration guide.

Then inspect confidence ranges separately: do more concentrated answers show better label accuracy in this dataset? That is an empirical question. Do not replace the event probability with confidence on the reliability diagram.

For a defined binary target, also report the Brier score, the mean of (p − y)², where y is 0 or 1. It evaluates probability predictions; it does not isolate calibration or measure trading returns. See the Brier score reference.

Keep sparse ranges visible. Nearby market observations can be correlated, especially when their forecast windows overlap. Use non-overlapping windows or time-block uncertainty estimates, and report exclusions. If fitting a calibration mapping, fit it on separate development data and evaluate it on untouched later data.

4. Account for execution and costs

Run an execution assessment with explicit assumptions. Include maker/taker fees, spread, slippage, latency, partial fills, cancellation and relevant network costs. State which items are modeled and which remain unmeasured. Avoid subtracting spread twice when fill prices already incorporate it.

A quote touching a displayed price does not establish queue priority or a fill. Compare favorable and adverse fill assumptions. Report completed trades, unfilled orders, gross and net results separately. State the denominator of any simulated win rate and how open positions are valued. Include loss magnitude and drawdown; many small wins can coexist with larger losses. These are proposed checks, not results from this site.

5. Use later data and simple baselines

Develop questions, criteria and evaluation rules on an earlier period. Freeze them before testing a later period. Leave a gap that prevents overlapping outcome windows from leaking across the boundary. Changing a threshold after viewing the test results turns that period into development data. The TimeSeriesSplit documentation explains chronological splits; market-window gaps still require explicit design.

Compare event probabilities against a constant event-frequency baseline estimated only from development data. Compare any execution experiment against predeclared simple rules and a no-order reference, using identical inputs, periods and costs. Report differences and uncertainty, including whether a higher-confidence subset has reduced coverage. A confidence filter is not evidence of improvement by itself.

6. Understand what the current demonstration can show

The verified Jev Trader source configures the dashboard backend as MODEL=mock and dry run. Its local rules combine momentum, order-book imbalance, trade flow and deterministic noise. They do not call Jev. Public market data and simulated fills cannot establish Jev accuracy or real account returns. The trading workflow guide explains the integration; the current decision adapter does not consume an official confidence field.

For an output-format exercise, open Ask Jev Trading, edit a teaching question and inspect the returned structure. It generates random examples, does not evaluate your text, and executes no trades. Its confidence illustration uses normalized entropy, not a verified TypeSafe formula. Use it to practice reading fields, not to build a calibration dataset.

Jev Trader is independently maintained and builds on Jarrod Watts's original project. This article describes a research protocol; neither the mock dashboard nor the learning tool demonstrates that Jev has passed it.