Experiments
How should you compare multiple AI trading bots?
Design separate, documented bot experiments without confusing personalities, trade count or a short winning streak with reliable evidence.
Compare bots as documented experiments with clear differences, fixed risk boundaries and comparable evidence. Multiple bots can help organize separate methods, but they do not guarantee diversification or a higher win rate. TradeTwin supports requests for separate private workspaces; each one needs verified provisioning, account configuration and plan consent before activation.
Define what the experiment is testing
Decide which difference you want to understand: an entry condition, a market, a time window or a decision rule. Record a baseline and make the alternative explicit. Changing sizing, entries and exits together makes it harder to interpret the result. A cautious or aggressive personality is an explanation style until it is expressed as reviewable rules. Give each proposed bot a clear purpose rather than judging a persuasive conversation as a trading experiment.
Separate workspaces without assuming separate risk
The intended private deployment model separates each bot's accounts, credentials and knowledge. That protects information boundaries; it does not make market risks independent. Two bots can react to the same gold move, share an exchange balance or hold correlated exposure. Review combined account capacity and total planned risk. A maximum number of positions in one bot cannot be interpreted as permission to multiply that number across several bots without considering the owner's overall limits.
Count the outcomes that belong to the test
Label the proposal version, relevant market data, entry-time reasoning, execution costs and exit outcome. Keep manual experiments distinct from outcomes attributed to the bot. If a transaction is excluded from learning, preserve the operational record needed for account reconciliation while ensuring the learning dataset respects that exclusion. A trade that was opened by the bot and manually closed needs that provenance recorded; otherwise a comparison can assign credit to a rule that was never followed.
Compare more than win rate
Win rate counts how often qualifying outcomes are wins, but it does not describe the size of wins and losses, drawdown, costs or the number of observations. A method can win more often and still lose money if losses are much larger. Use consistent inclusion rules and record missing data. A short series or a replay tuned to the same historical sample is a reason for further checking, not proof of a repeatable edge or a guaranteed future result.
Keep revisions deliberate
Use evidence to identify a proposed change, then compare the old and new plan before accepting the revision. New entry frequency should follow qualifying opportunities within saved limits, not a quota imposed on quiet market conditions. The public pilot does not automatically provision or activate requested bots. Independent private readiness and exact consent are necessary for every experiment. A documented inactive experiment is more useful than several active bots whose actual differences cannot be explained.
Questions worth asking
Do more bots automatically mean more diversification?
No. Bots may hold the same or correlated exposure even when their data is isolated. Review account overlap and combined risk. Separate storage is a privacy boundary, not a statistical claim about returns.
Which bot should I call the winner?
Use a consistent evaluation that includes sample size, costs, loss size and drawdown alongside win rate. The observed best bot in one sample may not remain best later; preserve the conditions and limits of the comparison.
Can requesting bots start trades automatically?
No. Requests are directory records. Each private workspace, account, executable plan and protection setup needs verification and the owner's review before execution can be enabled.