Evidence-based investing research
Value Investing Strategy (Strategy Overview)
Allocations for July 2026 (Final)
Cash TLT LQD SPY
Momentum Investing Strategy (Strategy Overview)
Allocations for July 2026 (Preliminary)
1st ETF 2nd ETF 3rd ETF

Investing Expertise

Can analysts, experts and gurus really give you an investing/trading edge? Should you track the advice of as many as possible? Are there ways to tell good ones from bad ones? Recent research indicates that the average “expert” has little to offer individual investors/traders. Finding exceptional advisers is no easier than identifying outperforming stocks. Indiscriminately seeking the output of as many experts as possible is a waste of time. Learning what makes a good expert accurate is worthwhile.

Deep Value Stock Selections by AI Panel

Is the evolving set of artificial intelligence (AI) platforms based on large language models interesting for selection of stocks that are deeply undervalued? Are they monolithic, or diverse? As a simple exploration, we pose to each of Grok, ChatGPT, Claude, Perplexity and Gemini the following prompt regarding undervalued U.S. stocks:

Adopt the persona of a deep value investor seeking the most undervalued publicly listed U.S. stocks of any size. List the three stocks most attractive to you as medium-term holdings. Do not provide any explanations.

We then compare and contrast results from AI panel members. Using responses to the prompt as posed at the end of June 2026, we find that: Keep Reading

Active U.S. Funds Not Quite So Bad?

The 2024 S&P Indices Versus Active (SPIVA) Scorecard finds that a supermajority of active U.S. funds underperform respective benchmarks. Is that finding representative of investor experience? In the May 2026 draft of their paper entitled “How the SPIVA U.S. Scorecard Understates the Performance of Actively Managed Mutual Funds”, Martijn Cremers, Jon Fulkerson and Timothy Riley restate the question addressed by the SPIVA Scorecard.

  • The Scorecard asks: what percentage of active funds either do not survive the full horizon or underperform the respective category benchmark indexes?
  • The authors ask: what percentage of active fund assets underperform equivalent passive funds?

They therefore adjust the SPIVA methodology, as follows:

  1. Instead of treating a fund that exits the sample as an underperformer, they use the actual returns of such a fund until its exit.
  2. Instead of weighting all funds equally, even though most assets are in a few large funds, they consistently weight by fund assets.
  3. Instead of using hypothetical benchmarks (total return indexes), they use existing equivalent passive funds.

They then compare SIVA results, replicated SPIVA results and adjusted results. Using categories and total returns from the same fund database used to generate SPIVA reports, along with returns for matched S&P benchmark indexes and passive tracking funds, during 2005 through 2024, they find that: Keep Reading

Performance of Active U.S. Mutual Funds and ETFs

What do the latest S&P Indices Versus Active (SPIVA) Scorecards say about active management investing expertise? In the “SPIVA U.S. Year-End 2025” Scorecard, Anu Ganti, Davide Di Gioia, Nick Didio and Liam Flaherty review the 2025 performance of active mutual funds and exchange-traded funds (ETF) per the following SPIVA Scorecard principles:

  • Account for discontinued funds, thereby eliminating survivorship bias.
  • Compare fund performance to that of a benchmark index matched to the fund investment category.
  • When aggregating funds, consider both equal-weighted and asset-weighted average returns.
  • Monitor investment style consistency over time to account for fund style drift.
  • For a given fund, use only the share class with the most assets to avoid double-counting classes.
  • Exclude all index funds, leveraged/inverse funds and other index-linked products.

Using categories and total returns for active U.S. mutual funds and ETFs and their associated S&P indexes during 2001 through 2025, they find that: Keep Reading

Explaining Earnings Announcement Stock Returns Using LLMs

Can artificial intelligence (AI) in the form of large language models (LLM) improve upon existing methods to explain stock returns around earnings announcements? In the June 2026 version of their paper entitled “Assessing the Benefits of Optimized Agentic AI Systems for Asset Pricing”, Ralph Koijen and Bradford Levy employ a a real-time, out-of-sample benchmark for evaluating optimized LLMs while avoiding lookahead bias and market adaptation effects. The benchmark measures how well AI systems explain stock returns around earnings announcements using only information available at announcement time (especially the announcement text). Optimized means experimenting with LLM prompts to improve results, such as guiding LLMs to assess earnings relative to expectations. Using earnings call transcripts, analyst consensus earnings and associated daily returns for U.S. stocks during the fourth quarter of 2025 (1,849 earnings announcements), they find that: Keep Reading

The State of Active ETFs

How should investors think about active exchange-traded funds (ETF)? In their February 2026 paper entitled “The Fast-Growing Market of Active ETFs”, Rachel Li and Nadia Winn review the state of active ETFs. They consider only ETFs offered by investment companies registered under the Investment Company Act of 1940. They identify active versus passive ETFs based on fund manager responses on SEC forms. Using assets under management (AUM), holdings and expense ratio data for the specified ETFs from SEC filings and fund tracking data from Morningstar during 2020 through 2024, they find that:

Keep Reading

AI-simulated CFO Sentiment

Can large language models (LLM) simulate market-scale, real-time business sentiment from the Chief Financial Officer (CFO) perspective? In their June 2026 paper entitled “CFOs Meet LLMs”, John Graham, Campbell Harvey and Manish Jha use an LLM (GPT-5.4) to simulate CFOs of specific firms. Simulations involve two prompts for each of 6,075 actual responses to the quarterly Duke-Federal Reserve CFO Survey during 2002 through 2025:

  1. System prompt – establishes context, persona and analytical instructions, including instructions to collect from the web CFO interviews and analyst reports, price targets and consensus revenue estimates. Instructions include a temporal restriction to ensure that information collected was available before the matched survey response date.
  2. User prompt – provides firm profile (industry, revenue, headcount and geographic footprint), respondent history (prior CFO optimism scores when available and the LLM’s own prior forecasts) and the key survey question: “Rate your optimism about the overall U.S. economy on a scale from 0–100, with 0 being the least optimistic and 100 being the most optimistic.” Again, a temporal restriction imposes an information cutoff date matching the actual survey date.

They run each pair of prompts three times a few seconds apart and average outputs to suppress LLM statistical variation. Finally they match every LLM average output to the response of the actual CFO at the same firm in the same quarter. Using past CFO Survey results and associated firm data spanning 2002 through 2025, they find that: Keep Reading

Investment Manager Selections by AI Panel

Is the evolving set of artificial intelligence (AI) platforms based on large language models interesting as a recommender of investment managers/advisors? Are they monolithic, or diverse? As a simple exploration, we pose to each of Grok, ChatGPT, Claude, Perplexity and Gemini the following prompt regarding selection of an investment manager:

Assume you are a U.S. investor with one to five million dollars to invest for retirement. Name the top three investment managers/advisors you would consider to assist you. Do not provide any explanation.

We then compare and contrast results from AI panel members. Using responses to the prompt as posed in late June 2026, we find that: Keep Reading

Retirement Portfolio ETF Allocations by AI Panel

Is the evolving set of artificial intelligence (AI) platforms based on large language models interesting as retirement portfolio specification advisors? Are they monolithic, or diverse? As a simple exploration, we pose to each of Grok, ChatGPT, Claude, Perplexity and Gemini the following prompt regarding exchange-traded fund (ETF) selections and weights:

For a hypothetical U.S. investor with moderate risk tolerance at each of ages 30, 40, 50, 60, 70 and 80, please provide your unique view on the three ETFs, with annually rebalanced fixed weights, that the investor should hold in a retirement portfolio over the next 10 years. Do not provide any explanation.

We then compare and contrast results from AI panel members. Using responses to the prompt as posed in early June 2026, we find that: Keep Reading

Safe Haven ETF Picking by AI Panel

Is the evolving set of artificial intelligence (AI) platforms based on large language models interesting as safe haven selection advisors? Are they monolithic, or diverse? As a simple exploration, we pose to each of Grok, ChatGPT, Claude, Perplexity and Gemini the following prompt regarding 16 potential safe haven exchange-traded funds (ETF) from the list in “Best Safe Haven ETF?”:

Using all training and real-time data available to you, please provide your unique view of 16 ETFs that are potential safe havens during U.S. equity market crashes by ranking them from best safe haven to worst safe haven: XLU, TLT, IEF, SHY, BIL, LQD, AGG, TIP, VTIP, VNQ, GLD, SLV, DBC, USO, UUP, GBTC. Do not provide any explanations.

We then compare and contrast results from AI panel members. Using responses to the prompt as posed in early June 2026, we find that: Keep Reading

Implications of LLM Use in Casual Investment Research

Is the shift from keyword-based search engines to artificial intelligence (AI) as implemented with large language models (LLM) affecting typical investor behavior? In the May 2026 revision of their paper entitled “The Double-Edged Mind: How LLMs Expand Stock Market Participation Yet Strengthen Confirmation-Seeking”, Cara Damm, Kevin Bauer, Florian Hett and Loriana Pelizzon address this question via an online experiment that randomly assigns participants to groups of volunteers who have access to keyword-based search engines, an LLM-based chatbot (unlabeled Gemini 2.0 Flash) or no information filtering tools. They design the experiment with two stages:

  • Stage 1: With access only to the name of the investment, participants initially choose a standard exchange-traded fund (ETF), a matched Environmental-Social-Governance (ESG) ETF or risk-free cash.
  • Stage 2: They revisit their decisions after receiving access to their respective assigned information filtering tools.

At the end, participants who chose the cash alternative keep their initial investments, while those who chose an ETF receive their initial investments adjusted by a fund return. Using responses from 374 participants in the experiment, they find that:

Keep Reading

Research Finder

Search 1,200+ research articles