Lower risk? Not in this data, and not provable either way.
Two hundred and twenty UK equity funds. Nineteen thousand five hundred and seventy-seven fund-months. Five risk factors, a momentum factor, and every risk-adjusted measure the literature argues about. The hypothesis was that Shariah-compliant funds carry less risk; it failed. Then the sample turned out to be too small to have tested it, and that, worked out properly, is the result worth publishing.
Are Islamic mutual funds exposed to lower risk than conventional funds? Evidence from the United Kingdom–
Every univariate test returns p > 0.21. The Islamic dummy in the six-factor model is positive but insignificant (p = 0.164), and its confidence interval straddles zero. Only drawdown gets close (p = 0.068), and the power calculation on this page shows that comparison had roughly a 7% chance of finding the gap it was looking for.
The question, and why it is not obvious
Shariah-compliant funds cannot hold interest-bearing instruments, cannot hold conventional banks or insurers, cannot use leverage beyond a narrow ratio, and screen out alcohol, gambling, tobacco and adult entertainment. Two arguments follow, and they point in opposite directions.
Screening removes the risky part
Excluding highly leveraged firms strips out the credit-like channel that turns market falls into fund collapses. Excluding banks avoids the sector that led the 2008 drawdown. On this reading Islamic funds should show a lower beta, lower volatility, and shallower falls (Hayat & Kraeussl, 2011; Naveed et al., 2020).
Screening removes the diversification
Markowitz (1952) is unambiguous: a smaller investable universe means a worse efficient frontier. Cut out whole sectors and what remains is concentrated, heavier in technology and healthcare, more dependent on one market regime, carrying more fund-specific risk (Hoepner et al., 2011; Walkshäusl & Lobe, 2012).
Theory cannot settle it. That makes it an empirical question, and the UK is the natural place to ask it: London is the largest Western centre for Islamic finance, Shariah pension inflows hit records in 2025, and yet the peer-reviewed UK fund evidence amounts to essentially one paper (Reddy et al., 2017). Almost everything else in this literature comes from Malaysia, Pakistan or the Gulf.
Islamic funds exhibit lower overall risk in the UK: systematic, idiosyncratic and downside.
Stated so it could fail. It did.
What the literature already said, and why it disagrees with itself
Sixteen studies, tagged by what they concluded. The disagreement in this field is not noise. It sorts almost perfectly by three things: which market was studied, whether the metric was a return or a risk, and whether the window contained a crisis.
The pattern: emerging-market studies that measure risk tend to find Islamic funds lower-risk. Developed-market studies that measure return tend to find no difference or a small penalty. Index-level studies that split by regime find Islamic portfolios look best precisely when everything else looks worst. A study's conclusion is largely determined before it collects a single price, by where it looks and what it decides to call risk.
The gap this dissertation aims at: a UK sample, recent enough to include COVID and the 2022 rate shock, run through the modern five-factor framework rather than CAPM alone, and judged on risk measures rather than mean returns.
The data, and the problem sitting inside it
Three. Not three hundred, not thirty, three UK-domiciled Shariah-compliant equity funds with a usable NAV history over the window, against 217 conventional ones. This is not a sampling choice that could have been made differently; it is the size of the UK Shariah equity fund market. Every result below has to be read through this picture, and the power section makes the consequence precise instead of leaving it as a caveat.
–
A note on the Islamic series: its last three months are identical to five decimal places. The index stopped updating in the export after May 2025 and the sample runs to July. It is two months of a ninety-month window and it does not touch the fund panel, but a chart should say when its own data has gone flat rather than let a reader treat a stale line as a calm market.
Nineteen thousand seven hundred and ninety-seven monthly excess returns. On a linear count axis this looks like a tidy bell; on a log axis the tails appear, and they are enormous: excess kurtosis of 513 against the 0 a normal distribution would give. The gold curve is that normal distribution, fitted to the same mean and standard deviation, and it is wrong at both ends. Four observations sit outside ±50% in a month. Those are not markets; they are NAV series with corporate actions in them, and they are the same funds that show up as outliers in the volatility cross-section.
The distribution is also centred in the wrong place: its mean sits at −7.3% a month, not near zero. That is not a market fact and it is not a rounding error. It is worked through in reading my own numbers back, along with what it does and does not invalidate.
Method, stated so it could be repeated
Four layers, each answering a question the one before it cannot. Nothing here is exotic; the point of using the standard toolkit is that a reader can check it against the papers it comes from.
Excess returns
Monthly simple returns from NAV, minus the UK 3-month Treasury bill. Everything downstream is denominated in the return an investor earned above doing nothing.
ri,t − rf,t = (NAVi,t / NAVi,t−1 − 1) − rf,t
CAPM: beta, alpha, and what is left over
Regress each fund's excess return on the market's. The slope is systematic risk. The intercept is Jensen's alpha. The standard deviation of the residuals is idiosyncratic volatility, the part of a fund's movement the market cannot explain.
ri,t − rf,t = αi + βi(rm,t − rf,t) + εi,t
IVOLi = σ(εi) = √( Σε²i,t / (T − 2) )
Fama–French five-factor plus Carhart momentum
A single beta assumes the market is the only systematic risk worth paying for. It isn't. Size, value, profitability, investment and momentum absorb the style tilts that screening creates, so that whatever is left in the dummy is a difference between fund types, not a difference between growth and value portfolios wearing different labels.
ri,t − rf,t = α + β·MktRFt + s·SMBt + h·HMLt + r·RMWt + c·CMAt + m·MOMt + γ·Islamici + εi,t
γ is the entire research question. If Shariah screening changes what an investor gets, after paying for every exposure the literature knows how to price, it shows up there and nowhere else.
Drawdown
Volatility is symmetric and memoryless; investors are neither. Maximum drawdown measures the worst peak-to-trough fall actually lived through, which is the number that makes people sell (Magdon-Ismail & Atiya, 2004; Chekhlov et al., 2005).
DDi,t = NAVi,t / maxs≤t NAVi,s − 1
MDDi = mint DDi,t
Group comparisons use Welch's two-sample t-test: unequal variances, which is the right default when one group has three members and the other has 217. Fund-level measures are equally weighted across funds; the pooled regression is not, because an unbalanced panel weights each fund by how many months it contributed.
Result 1: the univariate comparison
Every gap points the way the hypothesis predicted: Islamic funds look slightly better on Sharpe, slightly less volatile, slightly better on Treynor and alpha. Not one of them survives contact with a t-test. Beta is the starkest: 0.9148 against 0.9149, a difference in the fourth decimal place, p = 0.9923. Whatever Shariah screening does to a UK equity fund, it does not change how much of the market the fund is carrying. Move the threshold to 10% and nothing changes, because the smallest p-value in the table is 0.21.
Result 2: six factors, three specifications
–
The market factor is left out of the plot above on purpose: at 0.918 it is two hundred times the size of everything else and would flatten the rest into a vertical line. It gets its own exhibit, because what it shows is worth seeing on its own.
0.918135 pooled. 0.917432 Islamic-only. 0.918144 conventional-only. Three regressions, run on samples that differ by a factor of seventy-three, agreeing to three decimal places. Whatever else screening does, a UK Shariah equity fund is a UK equity fund: it carries the same market risk, and the intervals overlap almost entirely.
What does differ is style. HML is negative everywhere: a growth tilt, exactly what removing leveraged financials from an index produces. CMA is more negative in the Islamic-only fit (−0.0043 against −0.0021), and momentum, significant in the conventional sample, vanishes entirely for Islamic funds (p = 0.899), consistent with a mandate that forbids the speculative trading momentum is harvested by. Different constraints, different style exposures, indistinguishable outcomes.
Switch to t-statistics and a second lesson appears. In the conventional-only model almost every factor clears ±1.96, not because those effects are large but because 19,313 observations will find significance in almost anything. In the Islamic-only model, on 264 observations, most of the same coefficients cannot be distinguished from zero even though they have the same signs and similar magnitudes. Sample size is doing more work in this table than economics is.
Result 3: idiosyncratic risk
If screening concentrates a portfolio, the concentration should show up as risk the market cannot explain: the scatter left over once beta has done its work. This is the measure most likely to catch a diversification penalty, and the one the "narrower universe" argument predicts most directly.
–
–
Result 4: the only result that came close
–
A 5.7 percentage point difference in the worst fall an investor had to sit through, in the direction the hypothesis predicted, at p = 0.068. At a 10% threshold this is a result. At 5% it is not. It would be easy to write it up either way, and plenty of papers would. The honest treatment is to work out what the test could have seen, which is the next section, and it changes the reading completely.
Note what the chart is and is not. Both curves are averages across funds, so they show the shared shape of the period (COVID in early 2020, the rate shock through 2022) with fund-level dispersion averaged out. That dispersion is precisely what the t-test needs in order to conclude anything, which is why a chart that looks like a clear separation and a test that returns p = 0.068 are both telling the truth about different things.
What a sample of three could ever have detected
This section is not in the submitted dissertation. It is the calculation I would insist on now, and it is the part of this work I would defend hardest, because it converts a limitation everybody writes in their final chapter into a number.
A statistical test that finds nothing has two possible readings: there was nothing there, or the test could not see. Distinguishing them takes a power calculation, asking, before looking at the answer, how large a difference this design would have caught. With three funds in one group, the answer is unforgiving.
–
There is a further twist, and it is the reason no amount of extra data collection would have rescued this design. Precision in a two-sample comparison is governed by the smaller group. Raising the conventional count from 217 to a million moves the detectable difference by well under a percentage point. The 217 funds are not the constraint and never were. Three is.
This study did not find that UK Islamic funds carry the same risk as conventional ones. It found that a sample of three cannot tell.
Those are different sentences, and only one of them is supported. Every conclusion on this page is written to be the second.
Reading my own numbers back
Building this page meant extracting the underlying series out of the submitted document and recomputing things. One number does not survive that, and burying it would be worse than the error.
| What was reported | Value | What it implies per month |
|---|---|---|
| Mean of the excess-return column | −0.0732 | −7.3% |
| Islamic Sharpe × its volatility | −2.0948 × 0.0356 | −7.5% |
| Islamic Treynor × its beta | −0.0758 × 0.9148 | −6.9% |
| Implied average excess return | – | ≈ −7% a month |
Three independent routes through the results table land on the same impossible number. UK equity funds did not lose seven per cent a month above cash for seven and a half years, over the same window the FTSE 100 rose. The level of the excess-return series carries a units artefact: a risk-free rate on one scale subtracted from returns on another.
What that does and does not damage is worth being precise about, because the instinct is to assume everything is ruined and the truth is narrower.
Anything that reads the level
- Sharpe ratios: a monthly figure of −2.09 is not interpretable and should not be quoted
- Treynor ratios, for the same reason
- The absolute size of Jensen's alpha
- The pooled intercept as a statement about abnormal return
Anything that reads a slope or a difference
- Market beta: a constant shift in both sides of a regression moves the intercept, not the slope
- Every factor loading, for the same reason
- The Islamic dummy: the offset is identical for both groups, so it cancels in the comparison
- Volatility, idiosyncratic volatility and drawdown, none of which touch the risk-free rate at all
- Every p-value in the study
So the research question survives intact and the presentation of the risk-adjusted ratios does not. The right correction is to rebuild the risk-free series on the same scale as the returns and re-derive Sharpe and Treynor; the ranking between groups would not move, because a common constant cannot reorder two groups.
This is on the page because a portfolio piece that only shows the parts that came out clean is a sales document. Finding this in your own submitted work, tracing exactly how far it propagates, and publishing the propagation map is a more useful demonstration of how somebody handles data than any result would be.
What it means
Different constraints, similar outcomes
In a developed market, a Shariah-compliant equity fund is an equity fund with a rulebook. It carries the same market risk, falls at the same time, and recovers on roughly the same schedule. The constraints show up in style (a growth tilt, no momentum exposure), not in the level of risk. That is consistent with Reddy et al. (2017), the one prior UK study, and inconsistent with the emerging-market findings that motivated the hypothesis.
Why the emerging-market results don't travel
Pakistan and Bangladesh have dozens of Shariah funds, shallower markets, and conventional peers with far more heterogeneous strategies. Both halves of that make differences easier to see: more Islamic funds to measure, and more dispersion to measure them against. Neither condition holds in the UK.
For an investor with a religious constraint
The practical finding is the useful one, and it is reassuring rather than exciting: on this evidence, screening does not appear to cost return, and it does not appear to buy safety either. The "cost of conscience" (Renneboog et al., 2008) is not visible in UK data. Neither is a defensive premium.
For anybody designing the next study
Do the power calculation first. If the answer says the design cannot resolve a difference smaller than the quantity being measured, the fix is a different design (pooled European domiciles, a matched-pair construction, or fund-level bootstrapping), not a larger control group and a hopeful t-test.
Limitations, ranked by how much they matter
What I would do differently, in order: run the power calculation before collecting anything; widen the universe to UK-available rather than UK-domiciled funds, and to European domiciles if that is still too thin; cluster the standard errors; and pre-register the regime split so that "Islamic funds are more defensive in a crisis" is a hypothesis tested rather than a pattern noticed.
References
–
Show the reference list
The numbers are the submitted ones
Every table is transcribed from the dissertation. Every series (the FTSE 100 line, the Islamic index, both drawdown curves, the 216 fund volatilities, all 19,797 excess returns) was extracted from the data caches Word stores inside the document's charts. Nothing was read off a picture, and nothing was regenerated to look better.
The statistics run in your browser
Welch's t-test, Student's t distribution, and the noncentral t behind every power figure are implemented from scratch in about three hundred lines. The power sliders are not lookups against a table; they solve for the detectable difference each time you move them.
Checked against things that already have answers
141 tests on the engine: log-gamma and the incomplete beta against closed forms, t-quantiles against printed tables, Cohen's classic sample sizes for 80% power, the noncentral t collapsing to the central one at zero, and every transcribed coefficient re-checked so that t equals b over its standard error and each interval agrees with its own p-value.
The p-value is reproduced, not repeated
The dissertation reports p = 0.877487 for the idiosyncratic risk comparison. This page recomputes it from the printed group means and gets 0.877441, with entirely different code. That agreement is the reason to trust the rest of the arithmetic here.
What could not be published, isn't
The fund-level NAV panel is Bloomberg's and cannot be redistributed, so it is not here. What is here is everything derived from it that can be: the summary statistics, the distributions, the group series. Where a number is missing, the page says so instead of standing in a plausible substitute.
Drawn, not plotted
Every chart is inline SVG written by hand, no plotting library, nothing fetched. They inherit the site's own colours, so they work in both themes, print in black and white, and add nothing to the page weight worth measuring.
Need a question answered with data, and answered honestly?
Empirical finance work: panel construction, factor models, risk measurement and the power analysis that tells you whether the design can answer the question at all. Delivered with the code, the assumptions written down, and the limitations stated before anyone has to ask.