AlphaMaven AlphaMaven
The AI chatbots invading the stock market
Back to News
AlphaMaven Alternative Investment News

The AI chatbots invading the stock market

telegraph
3 months ago
Many experts believe it’s a matter of time before tech starts beating Wall Street’s best

More alternative-investment intelligence like this

Daily briefings, AI news summaries, and the full research platform on 372,600+ decision-makers & 118,200+ firms — start your free 7-day trial, card required, cancel anytime.

Start your free 7-day trial

News Summary available

## KEY TAKEAWAYS - A recent Nof1 research lab test found that six of eight leading AI chatbot models lost money on US tech stock investments, with Anthropic's Claude Sonnet declining nearly 60% on a $10,000 initial position and Google's Gemini losing over $5,000. - ChatGPT and Elon Musk's Grok were the only two models to generate positive returns in the test, with ChatGPT producing approximately $900 in gains and Grok breaking even, raising questions about AI's stock-picking reliability across different models. - Retail and professional investors are increasingly using large language models (LLMs) like ChatGPT and Claude to generate investment ideas, reflecting growing adoption of AI as a financial decision-making tool despite unproven track records. - Academic theory suggests that consistent stock-picking outperformance is inherently random and unpredictable, as demonstrated by Burton Malkiel's 1973 monkey-dart-throwing study, which challenges the premise that AI systems can achieve sustainable market edges. - Industry proponents, including former Meta engineer Faizan Ahmad and his startup Rallies, argue that AI stock-picking capabilities remain in early stages and will likely improve to eventually compete with top Wall Street performers. ## DETAILED SUMMARY Artificial intelligence chatbots are increasingly being deployed by both amateur and professional investors as tools for stock selection, yet recent empirical testing raises significant questions about their efficacy as investment vehicles. A test conducted by US research lab Nof1 evaluated eight of the most widely-used AI language models on US technology stock investing, producing mixed results that underscore the nascent state of AI-driven equity selection. The performance disparity among leading models was stark. Anthropic's Claude Sonnet suffered the worst outcome, declining nearly 60% from an initial $10,000 investment. Google's Gemini also underperformed substantially, losing more than $5,000. Among the eight models tested, six generated losses overall. Only two models achieved positive returns: OpenAI's ChatGPT generated approximately $900 in gains, while Elon Musk's Grok roughly broke even. These results suggest significant variability in how different AI architectures analyze and interpret market opportunities. The deployment of large language models in stock-picking contradicts established financial theory. Burton Malkiel's seminal 1973 research demonstrated that randomly selected portfolios—famously represented by monkeys throwing darts at the Wall Street Journal—performed comparably to professional fund managers over time. This theory implies that consistent outperformance is inherently random and that no systematic edge exists in equity selection. The current performance of AI models testing this hypothesis provides limited evidence supporting superior algorithmic performance. Despite recent test results, industry advocates argue that AI stock-picking remains early in its developmental cycle. Faizan Ahmad, a former Meta engineer and co-founder of Rallies, a startup leveraging AI for equity selection assistance, represents the optimistic camp contending that LLM capabilities will improve materially over time. Proponents believe it is only a matter of time before advanced AI systems beat top-performing Wall Street investors, though current evidence does not yet support this thesis. The divergence between bullish expectations and cautious empirical results suggests that institutional investors should maintain skepticism pending more extensive performance data across diverse market conditions and time horizons.