Best Practices for Creating Private Fund Performance Benchmarks

Best Practices for Creating Private Fund Performance Benchmarks
5:58

 

Data sourced from Dakota Private Markets, the private fund performance platform powered by Dakota. Learn More | Request Access

In Dakota's US private equity data, the top-quartile net IRR threshold ranges from 11.5% for the 2005 vintage to 27.3% for the 2023 vintage (Dakota Private Markets, as of June 30, 2026). A quartile label means very different things depending on which funds sit behind it, and that makes the size of the peer group as important as its definition.

The push in benchmarking has been toward narrower peer groups, and for good reason. A software-focused middle market buyout fund should not be ranked against industrials funds, and a 2019 vintage should not be ranked against a 2023 vintage. But every filter you add removes funds from the comparison, and a peer group can get so precise that it stops being reliable.

This guide covers the best practices for creating a performance benchmark you can defend: building the right peer group, knowing when a sample is too thin, and presenting results allocators will trust.

Why Small Peer Groups Break Quartile Rankings

A quartile ranking splits a peer group into four equal slices, so the number of funds in each slice drives how stable the result is. In a peer group of 8, each quartile holds 2 funds. Beat one more fund and you jump a full quartile.

Peer group size

Funds per quartile

What one fund changes

4 funds

1

"Top quartile" means beating 3 funds

8 funds

2

One fund can move you a full quartile

12 funds

3

One late reporter can shift the median

20 funds

5

Cutoffs start to hold up to a single outlier

40+ funds

10+

Quartile bands reflect the cohort, not a few names

Thin samples cause three specific problems:

  • Outliers set the cutoffs. With 8 funds, the top-quartile threshold sits between the 2nd and 3rd best performers. One unusually strong or weak fund moves that line by several points of IRR.
  • Late reporters reshuffle the rankings. Funds report on different schedules. When two new funds enter a 10-fund cohort, every ranking in it can change even though no fund's performance did.
  • The median stops meaning anything. The median of 6 funds is the average of two funds. It describes those two managers, not a market.

Vintage alone already produces the swings shown above. Stack strategy, geography, and sector filters on top of a single vintage and the remaining sample can get small fast. For more on how quartile bands are built, see Why Quartile Rankings Matter in Private Fund Performance Evaluation.

How Many Funds Is Enough?

There is no single magic number, but the arithmetic gives a practical scale. Use it to decide how much weight a quartile claim can carry.

Peer group size

How to treat the result

Fewer than 10 funds

A list of comparables, not a benchmark. Name the funds and compare directly instead of quoting a quartile.

10 to 19 funds

Usable as directional context. State the sample size alongside any quartile claim.

20 to 39 funds

A credible peer group for most LP reporting and IC discussions.

40+ funds

Stable quartile bands that hold up to late reporters and outliers.

The right threshold also depends on how the result will be used. A GP checking its own positioning before a fundraise can work with a thinner sample than an LP writing a re-up recommendation for an investment committee. The higher the stakes of the decision, the larger the sample should be.

Whatever the size, disclose it. A quartile claim with the cohort definition and fund count stated next to it gives an allocator something to evaluate. A quartile claim with neither gives them a reason to discount it.

Five Best Practices for Building Your Benchmark

The goal is the narrowest peer group that still has enough funds to rank against. These five practices get you there.

  1. Add filters in order of importance. Start with asset class and strategy, then vintage, then geography, then sector. Check the fund count after each layer so you know which filter shrank the sample.

  2. Drop the least important filter first. If the count falls below your threshold, remove the filter that matters least for this fund's return drivers. For a software-focused buyout fund, sector usually matters more than a narrow geography.

  3. Widen the vintage window openly, not quietly. Moving from a single vintage to a three-year window like 2019 to 2021 is a reasonable way to rebuild a sample. Label it as a blended window so no one mistakes it for a single-vintage benchmark. Peer Group Construction: Why Your Benchmark Might Be Comparing Apples to Oranges covers this tradeoff in more detail.

  4. Run a broad and a narrow benchmark side by side. If a fund is second quartile in the broad strategy benchmark and first quartile in a sector-specific one, show both. The gap is useful information, and showing both is more credible than picking the flattering one.

  5. Read all the metrics, not just IRR. In a small sample, a fund can rank differently on net IRR, TVPI, and DPI. Consistent rankings across all three carry more weight than a single strong number. For a refresher on each metric, see How to Evaluate Private Fund Performance in 2026: IRR, TVPI, DPI Explained.

Build a Benchmark You Can Defend

Dakota Private Markets covers performance for 18,000+ private funds across seven asset classes. Filter by strategy, vintage year, geography, fund size, and portfolio company sector, and see how many funds match your peer group criteria before you rank against them.

Request access to Dakota Private Markets and benchmark fund performance with data built for the comparison you're trying to make.

Peter Harris, Investment Research Associate

Written By: Peter Harris, Investment Research Associate

The Database For Cold Outreach to Reach Institutional and RIA Investors

The Database For Cold Outreach to Reach Institutional and RIA Investors