The Utah Review of Hormonal Treatments for Gender Dysphoric Minors: Methodological Appraisal
Downloads
In May 2025, the Utah Department of Health and Human Services (DHHS) published an evidence review of hormonal interventions for gender-dysphoric minors.
In May 2025, the Utah Department of Health and Human Services (DHHS) released an evidence review of hormonal interventions for minors. This analysis, which has been colloquially termed the “Utah Review,” concluded that hormonal interventions for gender-dysphoric youth are safe and effective. Proponents of pediatric gender transition have subsequently referenced the Utah Review as a counterpoint to the UK’s Cass Review and the U.S. HHS Review, which examined similar evidence but came to the conclusion that the benefits of hormonal treatments are uncertain, while the risks range from uncertain to certain and significant (e.g., harms to fertility).
Unfortunately, the Utah Review substituted quantity for quality and produced more than 1,000 pages of content that is deeply methodologically flawed, poorly organized, and frequently incoherent. The methodology used by the Utah research team contravened the standards for trustworthiness established by the National Academies of Sciences, Cochrane, and other expert bodies in evidence-based medicine. The Utah Review omits one of the most foundational aspects of a systematic review—a formal synthesis of evidence assessing its certainty. The reasons for this fundamental omission are not transparently disclosed, with contradictory explanations throughout the Review. Unencumbered by the methodological requirement to recognize the quality and certainty of the evidence, the Utah Review inappropriately concluded that pediatric gender transitions have significant benefits.
The problems with the Utah Review extend beyond its failure to recognize that its claims of proven benefits are methodologically untrustworthy. The analysis of harms is even more problematic. The researchers chose to deprioritize the analysis of outcomes of fertility and desistance—a decision made early in the process that biased the rest of the analysis, and ultimately misled researchers to conclude that there are no harms. They also paradoxically chose not to incorporate the findings of elevated mortality and morbidity among hormonally-treated individuals found in long-term studies into the Review's conclusions, but instead sequestered long-term results in a separate analysis (Part II, "Long-Term Outcomes). This decision signals a failure to recognize that concerns about problematic long-term outcomes of hormonal treatments for gender dysphoria are at the very core of the current debates about this treatment model for youth.
Numerous methodologically-questionable or indefensible decisions compounded throughout the analysis, culminating in conclusions of net-benefit for pediatric transition that are not only unsupported by evidence, but also contradicted by every credible systematic review conducted to date.
The goal of our detailed analysis is to help readers understand the Utah Review's content—as well as gain deeper insight into the nature of the problems in its key analyses, which signal critical limitations regarding the utility of the Utah Review for evidence-based decision-making.
I. Summary
Background
In May 2025, the Utah Department of Health and Human Services (DHHS) published an evidence review of hormonal interventions for gender-dysphoric minors. The evidence review was accompanied by a recommendation report regarding the provision of treatments. This work was commissioned by the Utah Legislature in 2023, following a moratorium on new hormonal interventions for minors diagnosed with gender dysphoria.
The bill (S.B. 16) stipulated that the moratorium could be lifted if a systematic medical evidence review of hormonal transgender treatments found that providing puberty blockers and cross-sex hormones for minors with gender dysphoria was a beneficial practice. The completion of this work (the “Utah Review”) was delegated to the Utah DHHS. The Utah DHHS commissioned the evidence review (“Evidence Review”) from the University of Utah College of Pharmacy’s Drug Regimen Review Center research team. Upon completion of the Evidence Review, DHHS prepared a recommendation report (“Recommendation Report”).
In May 2025, both the Evidence Review and the Recommendation Report were completed and published on the Utah State Legislature website.
Evidence Review
The Evidence Review was completed by the University of Utah College of Pharmacy’s Drug Regimen Review Center (DRRC) research team, and submitted by the director of the Center, Joanne Lafleur. Among the numerous analyses contained within the 1,051-page, two-part document, the analysis most germane to the Utah Legislature’s request for a systematic review of the evidence on the effects of puberty blockers and cross-sex hormones is the review of clinical studies. Unfortunately, this central analysis was not properly conducted because the researchers omitted a key step—a formal evidence synthesis assessing the evidence for certainty.
Despite this major omission, the DRRC research team proceeded to issue conclusions about the “consensus of the evidence” regarding puberty blockers and cross-sex hormones. They concluded that these interventions “are effective in terms of mental health [and] psychosocial outcomes” and “are safe in terms of changes to bone density, cardiovascular risk factors, metabolic changes, and cancer” (p. 90).
The omission of a formal evidence synthesis methodology is a major methodological violation. It contravenes the systematic review requirements set out by the National Academies of Sciences, Engineering, and Medicine—and indeed every published methodology for conducting systematic reviews.
The Utah Review offers contradictory explanations for this critical omission. The Evidence Review itself asserts that the researchers “were not contracted to include a synthesis of the evidence” (p. 90). Yet the accompanying Recommendation Report by the Utah DHHS provides a different explanation, stating that during the interim presentation of the results, it became apparent that there were “insufficient resources for DRRC to conduct a full evidence synthesis” (p. 4, Recommendation Report).
Regardless of the reason for the omission, without a formal evidence synthesis, the analysis conducted by the DRRC team cannot be considered a trustworthy systematic review—and in fact, it cannot be considered a systematic review at all. However, because the DHHS Recommendation Report accompanying the Evidence Review described the work as a “systematic review,” we applied two well-recognized systematic reviews analysis tools to assess its quality: ROBIS (Risk Of Bias In Systematic Reviews) and AMSTAR 2 (A MeaSurement Tool to Assess Systematic Reviews). These analyses yielded a rating of "high risk of bias" and “critically low” confidence respectively—the lowest ratings possible.
Irrespective of whether the Utah evidence review should be viewed as a very poor-quality systematic review or as a low-quality narrative review, the conclusion is the same: the findings of this analysis are not trustworthy. Below, we highlight some of the key limitations of the Utah evidence review of clinical studies:
- Cherry-picking outcomes. Despite the Utah Legislature clearly articulating that it was interested in the identification of harms, such as infertility, and specifically requesting that the analysis focus on detransition (p. 83), the Utah researchers chose not to consider either as a “high priority” outcome. Paradoxically, fertility was excluded on the grounds that hormone-related fertility harms are already known (p. 4, Recommendation Report). Since "high-priority" outcomes were used to screen studies for eligibility, this decision had multiple downstream effects, ultimately leading the researchers to overlook important signals of harm.
- Failure to conduct a systematic evidence synthesis. A trustworthy systematic evidence review is more than a collection of analyzed studies—it requires a structured, formal evidence synthesis and an evaluation of the certainty of the summarized evidence and the interpretation of its findings. These core elements are missing from the analysis. The researchers did not synthesize findings at each outcome level or assess the body of evidence for certainty. This critically compromises the evidence review and renders its conclusions untrustworthy.
- Failure to incorporate long-term outcomes. A long-term outcomes analysis, presented in Part II of the document, overlooked some of the most robust long-term outcome studies. However, it still found “increased mortality risks in both transgender men and women treated with hormonal therapy” (p. 914) but these findings were not incorporated into the main conclusions, which state that hormonal interventions are “safe in terms of changes to bone density, cardiovascular risk factors, metabolic changes, and cancer” (p. 90). As an outcome, mortality is among the most consequential, if not the most consequential, endpoints in clinical research. A signal of increased long-term mortality risk warrants explicit integration into the report’s principal conclusions. The failure to do so not only reinforces concerns about cherry-picking, but also reflects a lack of methodological rigor in the synthesis of the evidence.
Concealment of study details prevents meaningful engagement. Much of the 1,051-page report consists of tables of studies with accompanying text that has been redacted, purportedly to “protect” the identities of patients (see Figure 1 below). This highly unusual approach to concealing basic study information makes this section of the document difficult or impossible to assess. It is also nonsensical, since the referenced studies serve as the foundation for the practice of pediatric "gender-affirming" care and are already in the public domain, and in any case published academic studies generally do not reveal personally identifiable health information.
Figure 1

- Poor organization and usability. More generally, the presentation of the data makes it very difficult to understand, interpret, or verify. For example, the final list of 90 studies subjected to risk-of-bias analysis (pp. 54–62) appears in Appendix I.F without the reference numbers used in the main text and tables, and also includes abstracts, needlessly and substantially lengthening the section and rendering it impossible to efficiently extract crucial information. As a result, assembling a bibliography of these studies—or of those addressing the relevant clinical outcomes (pp. 44–53) of persistence, desistance, and regret (pp. 83–89), and long-term outcomes (pp. 904–906)—requires cross-referencing across multiple lists and tables. We ultimately compiled our own deduplicated lists, which we provide as supplements to this analysis.
- High number of errors. The haphazard presentation of data is further compounded by a high number of errors. For example, although the Review claims to have initially analyzed 134 primary studies (p. 44) the respective table and appendix list 230 studies (pp. 489–515). Table I.10 is described as listing 59 studies; in fact it lists 60 (p. 50). In Part II, the conclusion states that data were extracted from 15 studies out of a selected set of 25 (p. 914); in fact, the correct figures are 17 studies selected from 27 (p. 903). These and similar discrepancies throughout the text give the impression of a document assembled under time pressure and marked by unprofessional carelessness. There are also numerous data-extraction errors across the 90 final studies (e.g., Nahata et al., 2017; van de Grift et al., 2020).
In addition to the systematic review of the effects of hormonal interventions on gender-dysphoric minors required by the Utah Legislature, the bill also requested several other analyses. Those analyses were either not completed at all or were addressed at such a superficial level that they cannot reasonably be considered substantive responses, including:
- Inadequate handling of desistance-related research. Key detransition studies were overlooked as a direct result of the researchers' decision not to consider desistance a “high priority outcome” (p. 83), which negatively impacted the quality of the study search and analysis. The decision to de-prioritize desistance and related phenomena such as detransition and regret is inexplicable, given the Utah Legislature's explicit requirement of this analysis.
- Failure to include analysis of the effects of interrupting normal puberty. This analysis was requested by the Utah legislature but there is no sign that it was conducted.
While the University of Utah DRRC researchers did not complete the required analyses properly (or at all), they undertook several additional analyses that were not requested. These, too, were not properly conducted:
- Flawed overview of systematic reviews. The researchers failed to conduct an evidence synthesis, and instead largely confined the analytic exercise to critiquing the Swedish systematic review that underpinned Sweden's policy shift toward sharply limiting pediatric transitions. The researchers also made no mention of the NICE systematic reviews—issued in 2020—or the University of York reviews, published nearly a year before the Utah Review was made public, both of which informed the UK’s decision to restrict medical transition for minors. An overview of systematic reviews requires comprehensive and transparent efforts to identify all relevant reviews, including both peer-reviewed publications and relevant reports from reputable sources such as government agencies.
- Inadequate summary of clinical guidelines. As part of their response to a systematic review, the University of Utah researchers provided an analysis of clinical guidance documents. However, the analysis identified only 5 relevant documents—compared to the York systematic review’s assessment of 23 such guidelines and guidance documents. Despite applying very broad definitions of what constitutes a "clinical practice guideline," the University of Utah researchers' analysis overlooked some of the most influential clinical guidance documents shaping clinical practice in the English-speaking world— such as the American Academy of Pediatrics Policy Statement on care in the United States; the Australian Standards of Care and Treatment Guidelines, influencing practice in Australia; and the Cass Review in the UK. Nor were the current Swedish of Finnish national guidelines acknowledged by the analysis (however, the researchers did identify the outdated S1 German guideline as relevant). Much of the analysis focused on WPATH and the Endocrine Society's recommendations, but rather than assessing them for quality using the appropriate tools, the researchers simply asserted that because these documents are authored by "recognized medical authorities," they are by definition "evidence-based."
In fact, of the multiple analyses attempted by the University of Utah researchers, only one appears to have been adequately completed—the review of pharmacological agents used to treat gender dysphoria in youth. The authors identified 66 drugs in the category of puberty blockers or cross-sex hormones and determined that none are FDA approved for use in the treatment of gender dysphoria. Even with that analysis, however, the University of Utah researchers heavily editorialized its finding of the off-label status of the drugs used, strongly suggesting that this practice is universal and uncontroversial, and that it therefore should be immune from scrutiny or risk/benefit analysis.
Recommendation Report
The provenance of the Recommendation Report, which accompanies the Evidence Review, is less clear. Although the report was issued by the Utah DHHS, it appears that the University of Utah DRRC research team may have provided a draft which was reviewed by some DHHS advisers.
The Recommendation Report summarized the Evidence Review’s findings and issued recommendations related to the provision of care to gender-dysphoric minors. Although the Report recognized some methodological limitations of the Evidence Review, it continued to mistakenly refer to it as a “systematic review” and accepted the unfounded conclusion that puberty blockers and cross-sex hormones are safe and effective treatments for gender-dysphoric youth.
As a result of its uncritical acceptance of the Evidence Review, the Report focused on how to enhance the delivery of hormonal interventions for minors. The Report’s four main recommendations were to:
- Establish a hormonal transgender treatment board managed by DHHS that creates a “high quality certification and education” program for hormone prescribers;
- Restrict prescribing authority to “demonstrated experts”;
- Require a “comprehensive interdisciplinary care team model” that evaluates and reports outcomes;
- Institute an enhanced informed consent and assent process.
The Recommendation Report failed to discuss the circumstances under which these treatments should be contraindicated, despite having been explicitly instructed to do so by the Utah legislature. Notably, none of the recommendations addressed alternative or supportive treatments, such as specialized psychological care for affected minors.
Conclusion
The Utah Review’s central conclusion that hormonal treatments for gender-dysphoric minors are effective and safe was derived through a highly irregular evidence review process that omitted a fundamental step in systematic review development: a formal and transparent evidence synthesis assessing the certainty of the body of evidence. The term “evidence-based” is often misused as a simple synonym for “supported by some research,” or for any conclusion that follows a literature review. However, within the field of evidence-based medicine, “evidence-based” has a narrow, technical meaning. A review can only be considered genuinely evidence-based if it fulfills the requisite steps in a structured, transparent, and systematic manner. The Utah Review does not meet these criteria.
Although the political nature of the Utah Review’s provenance does not automatically invalidate the work, it does warrant heightened scrutiny of the Review's scientific objectivity. The Utah Review was commissioned in the wake of legislation imposing a moratorium on initiating new adolescent medical gender transitions, and a published commentary preceding Review's publication acknowledges that the commissioning of the Utah Review was the result of “coordinated effort across advocacy organizations working with the bill sponsors” with the aim of "potentially reopening a limited pathway for adolescent Utahns to receive care in Utah.” In that context, the Review should meet a high bar for neutrality and methodological rigor. Instead, it falls short: it shows a recurring pattern of interpretive bias (“spin”) in how it frames evidence, characterizes studies, and draws policy conclusions—choices that repeatedly appear to support a pathway toward lifting the moratorium rather than neutrally assessing certainty in the evidence base for pediatric hormonal interventions.
Consequently, healthcare decision-makers who value adherence to the principles of evidence-based medicine cannot rely on the Utah Review’s characterization of the evidence, and should instead be guided by extant rigorous and comprehensive systematic reviews, evidence appraisals, and policy documents grounded in high-quality systematic review methodology. These include systematic reviews conducted in Europe and North America, as well as broader overviews of the pediatric gender medicine practice based on systematic reviews, such as the UK Cass Review, and the U.S. HHS Evidence Review. These analyses found the evidence of benefits to be uncertain and the evidence of harms to range from uncertain to certain and significant.
The inescapable finding that the evidence that gave rise to the practice of pediatric transition—and continues to support this practice—is deeply flawed and that risks likely outweigh benefits at the population level—can plausibly support a range of policy choices, from restriction to exceptional cases, to limitation to ethical and scientifically valid research protocols, and inclusive of decommissioning these interventions for minors from routine practice entirely on the basis of an unfavorable risk–benefit ratio. Societal resources should be directed toward work that can help societies to chart a responsbile path forward through rigorous analysis of long-term patient data, strengthened outcome surveillance and follow-up, high-quality primary research where ethical and feasible, and transparent ethical debate about autonomy and safeguards for interventions that permanently alter the physical development of otherwise healthy adolescents in distress.
By contrast, continued investment of socieatal resouce in biased research—with the Utah Review serving as the most recent example—along with what appear to international efforts by WPATH-aligned practitioners to produce non-evidence based “consensus guidelines" that present the practice of pediatric gender transitions as uncontroversial and beneficial—are a disservice to the vulnerable youth such efforts claim to protect, and society at large.
II. Detailed Analysis of the Utah Review
Below we provide a more detailed analysis of the Utah Review. The goal of our analysis is to apply established principles of evidence-based medicine (EBM) to assess the quality of Utah’s Evidence Review, as well as the validity of the resulting recommendations contained in the Recommendation Report. We present our analysis in 12 parts:
- Introduction
- Protocol and Deviations
- Identification of Pharmacological Agents
- Systematic Review of Clinical Studies
- Overview of Systematic Reviews
- Long-Term Outcomes
- Clinical Practice Guidelines
- Desistance and Regret
- Interruption of Normally Timed Puberty
- Recommendation Report
- Evidence of Interpretive Bias and Spin
- Conclusion
1. Introduction
In January 2023, the Utah Legislature instituted a moratorium on prescribing hormonal treatments to minors diagnosed with gender dysphoria. The bill, S.B. 16, “Transgender Medical Treatments and Procedures Amendments,” specified that the moratorium could be lifted if a “systematic medical evidence review of hormonal transgender treatments” demonstrated the treatment’s safety and efficacy. The stated goal of the commissioned analysis was to help the Utah Legislature determine whether to lift the moratorium. The Legislature allocated over $100,000 to this initiative and tasked the Utah Department of Health and Human Services (DHHS) with overseeing the process.
The bill’s numerous provisions regarding the requested research guidance can be reasonably grouped into two main types of outputs:
- Evidence Review. The bill called for a systematic evidence review of short- and long-term impacts of puberty blockers and cross-sex hormones, including an assessment of the quality of the evidence. It also requested several additional analyses, such as an assessment of harms from interrupting natural puberty and estimates of desistance rates and timing.
- Recommendation Report. The bill requested recommendations regarding the treatments, including identifying situations in which the endocrine interventions should not be provided; the information minors and parents should receive before consenting to hormonal treatments; best practices for how that information should be provided; and a description of the “assumptions and value determinations used to reach a recommendation.”
Evidence of poor recommendations
The Evidence Review was requested by the Utah Legislature from the Utah DHHS and was subsequently commissioned from the University of Utah College of Pharmacy’s Drug Regimen Review Center research team (DRRC). The Evidence Review was conducted over 16 months (April 2023–August 2024) and published in its final form in May 2025. The final document is a two-part, 1,051-page report titled Gender-Affirming Medical Treatments For Pediatric Patients With Gender Dysphoria comprised of multiple research outputs, some of which are described as “deliverables.” A brief overview of these outputs and their status is provided below:
- Identification of pharmacological agents used and their FDA licensure status (“first deliverable”, Part I). This analysis was requested by the Utah Legislature. The Review identified 66 pharmacological agents and determined that none are FDA-approved for the indication of gender dysphoria. This component appears to have been adequately completed.
- Summary of clinical practice guidelines for treatment of pediatric gender dysphoria (“second deliverable”, Part I). This analysis does not appear to have been requested by the Utah legislature. The researchers assessed only four guidelines, omitting influential English-language guidance documents such as the American Academy of Pediatrics Policy Statement. This component was not adequately completed.
- Systematic reviews and/or meta-analyses (SRMAs) (“third deliverable”, Part I). This analysis was requested by the Utah legislature. The researchers attempted two different types of systematic review—(a) systematic review of clinical studies and (b) an umbrella review of existing systematic reviews. There are serious problems at every step in the development of both reviews, culminating with failures to assess the body of evidence for certainty. This component was not adequately completed.
- Supplementary data (Part I). Approximately 95% of the 1,051-page Evidence Review consists of tables and study lists produced during the analytic process supporting deliverables 1–3. Much of the material has limited utility, compounded by the unusual choice to redact crucial information about the studies comprising the body of evidence assessed. This presentation impedes the comprehension and usability of the report.
- Long-term outcomes (Part II). This analysis was conducted as an add-on, after the interim presentation of results revealed a lack of the long-term outcomes analysis explicitly requested by the Legislature. While some findings were presented in Part II, the conclusion of elevated mortality in hormonally-treated individuals was not incorporated into the overall conclusions about intervention safety. This component was not adequately completed.
Following the completion of the Evidence Review, Utah' DHHS completed a Recommendation Report. In May 2025, both documents were published on the Utah State Legislature website.
2. Protocol and Deviations
The Utah Review describes itself as a "systematic review" of evidence. A trustworthy systematic review must be guided by a protocol— a predefined, explicit methodology to find and synthesize all the relevant evidence—in order to minimize bias and ensure transparency and reproducibility. While systematic review types vary based on research goals, all share the following required steps:
- Determining the study question and eligibility criteria
- Searching for evidence and selecting studies
- Abstracting data and assessing risk of bias
- Creating a formal evidence synthesis (i.e., summarizing the findings from the studies for each outcome of interest)
- Assessing the certainty of the evidence and drawing conclusions
It is not uncommon during the research process to identify barriers to completing the protocol as previously planned. When that happens, protocols should be modified, and any deviations must be explained and be methodologically appropriate.
The Utah Review's approach appears to have followed the following steps. [insert a 2-4 sentence description of how it is a single protocol driven by identification of pharmacological agents... then how it was adjusted further for each analysis)]. However, the Utah Review does not provide a pre-defined research protocol. Although the researchers assert that pre-specified “internal protocols” existed (p. 12), the Review describes only what was ultimately implemented—leaving readers unable to distinguish what was planned from what was done, or to assess whether any deviations were justified. Nonetheless, there is ample indication of substantial departures that may not be methodologically appropriate.
Key Deviations
Below, we outline key elements of a typical research protocol and summarize what can be inferred about the original plan and the deviations that appear to have occurred.
- Lack of pre-defined protocol. Protocol preregistration is considered best practice in evidence reviews to promote transparency, prevent selective reporting, and reduce bias by documenting methods and outcomes in advance. This ensures the review follows a predefined plan, making it easier to detect deviations or outcome switching, and also helps avoid duplication while enhancing credibility. However, when protocol registration is not possible or not done, it is acceptable for researchers to explicitly explain their rationale, and provide a detailed description of their methodology, including modification to the methodology during the review process to facilitate transparency. This Utah Review did not register their protocol, provide details regarding their methodology approach, or explain their apparent deviations from conventional methodological approaches—such as their notable omission of the evidence synthesis and assessment of the evidence quality/certainty.
- Unclear rationale for research methodology. Two different methodologies are consistent with the Utah Legislature’s requirement of a “systematic review” of studies evaluating outcomes of interventions. The first is a systematic review of clinical studies evaluating the effects of puberty blockers and cross-sex hormones. Such reviews have been completed by the UK’s NICE, the University of York, and McMaster University. The second is an overview of existing systematic reviews (also known as an “umbrella review”) on the same topic. This approach was used in the analysis commissioned by the Florida Agency for Health Care Administration and by the U.S. HHS evidence review. Depending on research goals and resources, research teams typically conduct either a systematic review of clinical studies or an overview of systematic reviews, but rarely both. The University of Utah team chose to attempt both approaches. Unfortunately, neither analysis was performed to an acceptable methodological standard. The Utah Review also attempted to complete a guidelines analysis and a separate long-term outcomes research. Neither was completed to a high methodological standard and for unclear reasons, long-term outcomes—specifically, signals of elevated morbidity and mortality— were not incorporated into the main review's conclusions.
- Lack of clearly defined research question. A credible systematic review process begins with articulating research questions, typically structured using the "PICO" framework (Population, Intervention, Comparator, and Outcome). In the Utah Review, only the population is explicitly defined—namely, pediatric patients (<18 years) described as experiencing gender dysphoria or identifying as transgender or non-binary (Table I.1, p.13). The remaining PICO elements are not transparently specified upfront. Instead, the interventions must be inferred indirectly from the search strategy, where specific endocrine interventions of puberty blockers and cross-sex hormones appear among numerous technical search terms (p.16). The comparators and outcomes can only be reconstructed retrospectively by examining eligibility criteria and study assessments (e.g., p.24). This lack of explicit PICO formulation makes it unclear whether these elements were predefined or modified post hoc, limiting transparency and raising concerns about potential protocol deviations and selective framing of the evidence base.
- Limited scope of literature search. The scope of literature search was limited regarding its time horizon, databases, and language—notably choosing not to search the websites of international public health authorities, which house much of the relevant information that was the focus of the Review. An example of the limited search strategy is the Utah Review's failure to identify clinical care guidelines from Sweden and Finland. In fact, the Utah Review identified a total of 5—and retrieved data from only 4 clinical guidance documents. In comparison, the systematic review by Taylor et al., 2024 included 23 guidelines, with 12 guidelines published after 2018. As for the review of systematic reviews, the Utah Review completely missed a series of recent systematic review publications examining the outcomes of puberty blockers and cross-sex hormones, published several months ahead of the publication of the Utah Review. The limited nature of the search introduces bias and undermines the completeness and generalizability of the Review's findings.
- Unexplained exclusion of "regret" from "high-priority" outcomes. According to a presentation given to the Utah interim DHHS committee on June 19th, 2024, the outcome of "regret" was considered a "high-priority outcome" (see slide 7, reproduced below in Exhibit X). However, this outcome is notably absent from the list of "high-priority outcomes" in the Utah Review itself (p.24), which was completed and submitted to Utah DHHS on August 6, 2024, just a few weeks later. There is no explanation for this deviation. Instead, the authors simply state that regret was "not among our high-priority outcome categories for this review." It is unclear when and why this outcome was dropped from the list of "high-priority" outcomes, given the researchers' own acknowledgment that the topic of regret was "pointedly of interest to the legislature" (p. 83).
- Unclear fate of the outcome of "infertility." The Recommendation Report accompanying the Evidence Review indicates that infertility was considered at some stage of the Review, but ultimately not designated a “high-priority” outcome. The stated rationale is that “infertility is a known risk of CSHT [Cross-Sex Hormone Therapy] and was not an outcome of focus in the DRRC systematic review” (Recommendation Report, p. 4). This approach to defining critical outcomes is incoherent: under the Review’s methodology, outcomes not deemed “high priority” are not analyzed and therefore cannot meaningfully inform the conclusions. As a result, any conclusions about the safety of endocrine interventions are, from the outset, incomplete.
- Inconsistent and unclear use of comparators. The Utah Review's approach to "comparators" is obscure. According to the same presentation given to the Utah interim DHHS committee of June 19th, 2024, the planned comparators included "before and after" hormonal intervention, "transgender with and without hormone treatment," and "transgender compared to cisgender." The Review lists the latter two, but not the "before and after" comparison (pp. 24, 71, 901). In the analysis of the studies for risk of bias, the Review deviates further, treating any subgroup reports (e.g., males and female outcomes reported separately) as though the study had a "comparator" (e.g., risk of bias analysis of Chen et al., 2023 on p. 556)—even when the studies themselves explicitly acknowledge that no comparators were used, as happens in Chen et al.
- Unclear reason for omitting evidence synthesis and assessing the evidence for quality/certainty. Further, the decision not to conduct a formal evidence synthesis or assess certainty of evidence also appears to be a deviation from the original intention, as suggested by the Recommendation Report. Problematically, the Evidence Review itself contradicts this explanation and asserts that a formal evidence synthesis was never intended in the first place. Such inconsistency in explanations suggests a lack of transparency. Regardless, the omission of a formal evidence synthesis renders the work product, at best, a non-systematic narrative review—vulnerable to bias and not a reliable basis for evidence-based decision-making.
Exhibit X - [title and link]

Brief conclusion about the protocol
3. Identification of Pharmacological Agents
The University of Utah DRRC team reports using standard databases (e.g., Micromedex, UpToDate, the FDA Orange Book) to "identify a comprehensive list of all drug product hormones and hormonally active agents that are used in pediatric TGNB [transgender and nonbinary] patients in the United States." The researchers identified 66 U.S. drugs prescribed for patients with gender dysphoria. The analysis found that none of the drugs are FDA-approved for the indication of gender dysphoria.
The analysis, however, went beyond statement of fact and provided a narrative to frame the non-FDA approved use of puberty blockers and cross-sex hormones in dysphoric youth as a normative, uncontroversial practice. The analysis opened with the statement, “off-label use of medications is widely accepted and often becomes a standard of care” and emphasized high prevalence of off-label use with a range so broad as to be essentially uninformative ("worldwide estimates range from 3.2% to 95%”) (p. 6). The researchers provided further reassurance that "common medications used off label" include “antibiotics…beta blockers…psychiatric drugs”, adding that "one study showed no differences in the risk of adverse events” between off-label vs indicated uses.
Identified Drugs
The implied equivalence between routine pediatric off-label use of short antibiotic courses for infection and endocrine drugs with lifelong consequences and known risks in otherwise healthy children is invalid. Unlike antibiotics—rigorously studied in adults against a well-understood disease course—endocrine interventions for gender dysphoria have not been properly tested or shown effective in either pediatric or adult populations, making the comparison inadequate.
In addition, the statement that "one study showed no differences in the risk of adverse events" is misleading and contradicts the preponderance of evidence. Multiple well-designed studies demonstrate that off-label use in pediatrics is associated with worse health outcomes, with significantly increased risk of adverse events (risk/odds ratios of 1.67-2.25).
Further, the fact that other medications are prescribed off-label with weak evidence does not necessarily justify additional off-label uses but instead highlights a broader problem requiring closer scrutiny.
Summary. While the analysis provides a helpful compilation of hormonal interventions used in “gender-affirming” contexts, the editorializing suggests interpretive bias rather than an impartial synthesis. The goal appears to position the off-label use of endocrine interventions as low risk and routine, and to ameliorate concerns about the potentially unfavorable risk-benefit profile of using these off-label interventions on minors with no underlying physical pathology.
4. Systematic Review of Clinical Studies
“…after having spent many months searching for, reading, and evaluating the available literature, it was impossible for us to avoid drawing some high-level conclusions.” (p. 90)
Despite having omitted the key analytical steps required to arrive at a methodologically sound conclusion, the researchers nevertheless issued a conclusion that hormonal interventions for minors are effective and safe
“… the consensus of the evidence supports that the treatments are effective in terms of mental health, psychosocial outcomes, and the induction of body changes consistent with the affirmed gender in pediatric GD patients. The evidence also supports that the treatments are safe in terms of changes to bone density, cardiovascular risk factors, metabolic changes, and cancer.” (p. 90)
The attempted systematic review of clinical studies encompassed 90 studies. Unfortunately, the researchers omitted a key step of a systematic review—formally synthesizing the evidence and assessing its certainty (quality).
We applied ROBIS, a widely used structured tool for assessing risk of bias in systematic reviews. A "risk of bias" is a methodological term that assesses the degree to which the systematic review process (how it was designed, conducted, or interpreted) may have skewed its conclusions. ROBIS evaluates 4 key areas of review conduct (“domains”) as well as an overall "risk of bias" score. As shown in the ROBIS summary (see Table 1), serious methodological concerns were identified across all 4 domains—including the domain of "synthesis" of evidence, which is completely missing from the Utah analysis. The ROBIS assessment yields the overall rating of "high" risk of bias (RoB)—the lowest possible rating. For detailed analysis, the full ROBIS assessment is available here.
Table 1. ROBIS Assessment of the Utah Review of Clinical Studies
| ROBIS | Study Eligibility Criteria | Identification and Selection of Studies | Data Collection and Study Appraisal | Synthesis and Findings | Overall RoB |
| Utah Review: Clinical studies | High | High | High | High | High |
We also applied AMSTAR 2, another systematic review assessment tool, which focuses on the overall quality of evidence reviews. This is the appraisal tool the Utah Review itself utilized in assessing other systematic reviews (see Section ii below). AMSTAR 2 is designed to assess the methodological quality of systematic reviews of healthcare interventions and distinguishes several “critical domains,” violations of which warrant a rating of low or critically low confidence. When applied to the Utah Review, AMSTAR 2 yields an overall rating of "critically low" confidence—the lowest possible category (see Table 2). A finding of “critically low” indicates that the review has more than one critical flaw and should not be relied upon to provide an accurate and comprehensive summary of the available evidence. Our assessment identifies 4 critical flaws, including the absence of a pre-registered protocol or prospectively defined methods (critical domain 2), and failure to incorporate risk-of-bias assessments into the review’s conclusions (critical domain 13). For detailed analysis, the full AMSTAR 2 assessment is available here.
Table 2. AMSTAR 2 Assessment of the Utah Review of Clinical Studies
| 1. PICO Components in Research Questions and Inclusion Criteria | No |
| 2. Protocol and Deviations | No |
| 3. Study Design Selection | Yes |
| 4. Literature Search Strategy | No |
| 5. Study Selection in Duplicate | Yes |
| 6. Data Extraction in Duplicate | No |
| 7. Excluded Studies | Partial Yes |
| 8. Description of Included Studies | Partial Yes |
| 9. Risk of Bias (RoB) Assessment | Partial Yes |
| 10. Funding Sources of Included Studies | No |
| 11. Meta-Analysis Methods | N/A |
| 12. Impact of RoB on Meta-Analysis | N/A |
| 13. RoB Consideration in Interpretation | No |
| 14. Heterogeneity Explanation | No |
| 15. Publication Bias | No |
| 16. Conflicts of Interest | Yes |
Taken together, ROBIS and AMSTAR 2 point to the same conclusion: the Utah Review falls well short of accepted standards for systematic reviews. Its findings therefore cannot be regarded as methodologically robust and are at high risk of bias.
Below, we examine some of the most significant deficiencies in greater detail:
- Lack of a pre-specified protocol. A clearly articulated research protocol is essential for the transparency and reproducibility of systematic reviews. While the researchers assert that some pre-specified "internal protocols" existed for some of the steps they followed, these are not presented anywhere; instead, the researchers explain what they did but do not disclose what they had planned to do, or how or why their final actions deviated from their plans. The study eligibility criteria appear post hoc and there is contradictory information regarding whether the researchers never planned to synthesize the evidence and assess it for certainty or rather ran out of resources and abandoned the plan. (In contrast, see a proper protocol for a systematic review of the effects of puberty blockers by McMaster University).
- Questionable outcome selection. Outcome selection is a key step in a systematic review. Unfortunately, the Utah Evidence Review did not list the outcomes of interest in the PICO format, and the selection of “high-priority” outcomes appears arbitrary at best (see Figure 2). The outcome of infertility was excluded from “high-priority” outcomes on the grounds that hormone-related fertility harms are already known (p. 4, Recommendation Report)—a puzzling justification in the context of an evidence review. The outcome of desistance was de-prioritized with no explanation, despite the Legislature’s explicit request that desistance be evaluated as a key outcome of interest. The outcome of mortality was not considered in the main analysis at all and was only addressed in a separate analysis of long-term studies, where it was not integrated into the main findings.
Questionable study selection process. The researchers identified 662 studies, of which 227 were deemed “eligible;” of those 90 were selected for the final analysis (see PRISMA flowchart, p. 34). The 90 selected studies span heterogeneous populations, interventions, outcomes, and study designs, without a clearly articulated rationale for how their selection contributes to a coherent analytic framework. Below are several examples of the 90 selected studies:
- A study comparing neural (brain) responses to unfamiliar peer versus familiar caregiver voices among testosterone-treated versus untreated transgender-identified teen females.
- A study examining insurance coverage for puberty blockers and mental health comorbidities in patients treated at a single-site clinic.
- A study comparing different doses of the same pharmacological puberty-suppression agent (e.g., histrelin implants Vantas and SupprelinLA) in terms of hormonal suppression in adolescents at Tanner stages 2 and 3.
While collecting heterogenous studies related to hormonal interventions generates an impressively long list, it challenging to see such an approach can aid in answering a specific research question—which is the point of a systematic review.
Inadequate assessment of the studies’ risk of bias. A risk-of-bias (RoB) assessment of individual studies is a key component of a systematic review because it enables an assessment of a study’s methodology to determine if the study’s reported treatment effects are trustworthy. However, the Utah Review’s RoB analysis was not properly conducted. The researchers did not appear to closely engage with the studies and did not take into account their methodological limitations.
A review of the Utah Review’s RoB assessments of several well-known studies suggests a broader pattern of inflated ratings, particularly for observational studies that lack true control groups. Alarmingly, in multiple instances the researchers rate studies with no control groups as though they had control groups by treating within-study subgroup analyses as if they were external comparators. For example, if outcomes are reported separately for males and females, the analysis implicitly treats females as a “control” for males (or vice versa). This practice manufactures a comparator where none exists, artificially elevating perceived methodological strength. The Utah analysis of two influential studies reporting benefits of youth transition—but marked by serious methodological limitations not adequately recognized in the Utah Review—are described below.
- Chen et al., 2023: The University of Utah DRRC researchers assigned Chen et al., 2023 a near-perfect 8/9 score on the Newcastle-Ottawa Scale (NOS), despite its serious limitations. The researchers treated this study as having a comparator/control group and awarded full credit for comparability, stating that the comparator was “drawn from the same community as the exposed cohort” (p. 556.) However, Chen et al. explicitly acknowledged that the study lacked a comparison group: "Finally, our study lacked a comparison group, which limits its ability to establish causality” (Chen et al., 2023, p. 249). The study also failed to report most pre-specified outcomes, had approximately 30% unexplained attrition at 24 months (besides reporting that 2 youths died by suicide during treatment), and did not adequately control for important confounders. Independent assessments of this study range between 3–5 stars on the NOS, indicating moderate to high risk of bias.
- Tordoff et al., 2022: The researchers’ “moderate risk of bias” rating of Tordoff et al., 2022 (6 of 9) is similarly inflated (p. 570). Although the study included an untreated comparator group, that group shrank from 35 to just 7 subjects by the study’s end. By definition, this subgroup differed systematically from the treated cohort, as they were the only study subjects who did not initiate treatment despite remaining in care. It is difficult to reconcile how such a small and self-selected group could reasonably be considered comparable to the treated cohort. Moreover, only 64 of 169 eligible subjects (<40%) remained at 12-month follow-up, and follow-up duration was insufficient to assess longer-term outcomes. Of note, the York systematic review assigned this study an NOS score of 3.5, classifying it as at “high risk of bias”—a substantially more conservative appraisal.
Shortcomings in risk-of-bias analysis
Compounding these concerns, neither data extraction nor RoB assessment was performed in duplicate, as recommended by the Cochrane Handbook to reduce bias and risk of error.
- Absence of meaningful evidence synthesis. An evidence synthesis, which includes assessment of the evidence for certainty, is a cornerstone of a trustworthy systematic review. The Review does not apply any standardized framework (e.g., GRADE: Grading of Recommendations, Assessment, Development, and Evaluation) or even a narrative synthesis (such as the one provided by the York reviews) to assess the certainty or strength of the evidence across outcomes.
For example, when the authors state, “rates of depression and suicidal thoughts/self-harm tended to be lower among hormonally treated transgender youth compared to untreated transgender individuals,” (p. 77) they do not specify which treatments were assessed, how many studies contributed to this conclusion, how many participants were included, which specific outcome measures were used, or the magnitude of effect sizes. As a result, readers are given no basis for determining whether findings are based on robust comparative data or rather low-quality observational evidence, small samples, and high attrition studies.
The researchers provide conflicting explanations for the lack of a full evidence synthesis. In the Evidence Review the authors claim they "were not contracted to include a synthesis of the evidence that we found: only to assess ROB and provide evidence tables summarizing safety and efficacy findings.” ( p. 90). This claim is difficult to reconcile with the Legislature’s explicit request for a systematic review, since evidence synthesis is a core component of a systematic review. Adding to the confusion, the Recommendation Report states that the reason for a lack of a proper evidence synthesis is that the researchers ran out of resources; “the DRRC presented a summary of its draft report… it became clear there were insufficient resources for DRRC to conduct a full evidence synthesis” (p. 4, Recommendation Report). Incoherent presentation of results. Absent an evidence synthesis, much of the information about the evidence assessed has to be gathered through data tables, which comprise much of the 1,000+-page document. However, the authors’ highly unusual choice to redact information about the studies, purportedly to “protect” the identities of patients and clinical centers (see below), makes the document challenging to evaluate and validate.
Some examples of these superfluous and obfuscatory redactions are provided in Figure 3 below:
Figure 3

- High number of errors. Our attempt at validation revealed multiple errors and irregularities. While the Review claims that “there were N=134 primary clinical studies reporting findings in TGNB populations all over the world” (p. 44), the relevant table and appendix list 230 studies (pp. 489–515), as does the conclusion (p. 90). No list of these studies is presented; we have deduced these studies through triangulation and provide the list here. A 2018 Dutch study selected for data extraction is inexplicably missing from the Review’s list of relevant clinical studies. In Nahata et al. (2017), anxiety counts were misreported (p. 659). In van de Grift et al. (2020) follow-up age and several outcome counts were inaccurately recorded (pp. 726–729). Spot-checking suggests that such errors are not isolated. These errors are likely partly due to the fact that the researchers deviated from the established practice of having two researchers verify the extracted data.
- Reliance on non-standard terminology to compensate for omission of key evidence appraisal steps. The conclusions of the Utah Evidence Review state that “consensus of the evidence supports that the treatments are effective.” However, “consensus of the evidence” is not a recognized term in evidence-based medicine (EBM), and it lacks any methodological underpinning. From the EBM perspective, acceptable terminology refers to the “body” (Guyatt, 2011, p.1294) or “totality” (Guyatt, 2015, p. 16) of evidence and should follow systematic review methodology, culminating in an explicit assessment of certainty (e.g., via GRADE), which evaluates risk of bias (potential for systematic errors in study design or conduct), inconsistency (variation in results across studies), indirectness (how directly the evidence addresses the clinical question), imprecision (width of confidence intervals), and publication bias.
The authors omitted these critical steps and instead described their process as follows: “after having spent many months searching for, reading, and evaluating the available literature, it was impossible for us to avoid drawing some high-level conclusions.” This does not constitute a valid method for assessing a body of evidence. Introducing new terminology does not remedy the underlying methodological omission of an evidence synthesis that assesses the body of evidence for certainty.
Summary: The University of Utah’s attempt at a systematic review did not fulfill the essential function of such reviews. A systematic review is more than a compilation of tables and citations; its value lies in a transparent, reproducible evidence synthesis that evaluates the certainty of the evidence prior to drawing conclusions. That step was not undertaken. The reasons for this critical omission are unclear, but it is likely that a poorly defined search strategy which yielded an unwieldy body of studies of questionable relevance rendered the subsequent steps of the analysis too difficult to complete. This appears to have resulted in a high number of errors in assessments of individuals studies, and failure to complete the final step of assessing the evidence for certainty. As such, any conclusions about the “evidence” that emerged from the Utah analyses of studies cannot be considered reliable.
5. Overview of Systematic Reviews
The DRRC research team analyzed seven systematic reviews (such analysis/overview of systematic reviews is also known as an “umbrella review”). Just as with the systematic review of clinical studies (discussed in the prior section), the most serious deficiency of the attempted umbrella review is the absence of an evidence synthesis and an assessment of the certainty of the evidence.
We analyzed the DRRC overview of systematic reviews using the PRIOR checklist for umbrella reviews. Several crucial shortcomings are outlined below:
- Inadequate search strategy and failure to capture recent systematic reviews. The researchers overlooked major databases and systematic review repositories—including the Cochrane Database of Systematic Reviews (CDSR), PROSPERO, CINAHL, and Epistemonikos. They also failed to search relevant gray literature, such as systematic reviews posted by public health authorities, including the UK NICE evidence reviews (2020) on puberty blockers and cross-sex hormones.
In addition, the researchers did not update their search after the publication of the highly relevant York systematic reviews on puberty blockers and cross sex hormones. As a result, the DRRC analysis missed the UK systematic reviews that have been foundational to UK policy changes on the provision of these interventions. - Subjective systematic review selection. The DRRC used a broad definition of systematic review, considering any review that describes itself as “systematic” as such, regardless of the methods used. The DRRC team identified 38 systematic reviews, seven of which were designated “high priority” for evaluation. However, the selection criteria appear to be inconsistently applied. For example, Mahfouda et al. (2019)—a review paper that is not a systematic review and does not position itself as such—was included and appraised as though it met systematic review criteria. Of note, this paper is frequently cited in support of provision of endocrine interventions to minors, which may explain its inclusion.
- Questionable assessment of systematic review quality. The DRRC rated the 7 “high-priority” systematic reviews using AMSTAR 2 (p. 43). Unfortunately, there appears to be evidence of bias in the DRRC team’s handling of the Swedish review (Ludvigsson et al., 2023), which has been the basis for Sweden’s practice shifts away from gender transition of minors. Specifically, the researchers erroneously cite the exclusion of high-risk-of-bias studies from synthesis by the Swedish researchers as a major methodological violation—yet this is a standard approach recommended in the Cochrane Handbook. Another umbrella review, that published by the US HHS, also noted some limitations in the Swedish review, but judged it to be at low risk of bias and among the field’s higher-quality systematic reviews.
- Inadequate evidence synthesis. Instead of providing a summary of what the collective body of systematic reviews demonstrates regarding mental health outcomes, physical outcomes, or long-term harms, the DRRC discussion of systematic reviews focuses nearly exclusively on criticizing the Swedish review by Ludvigsson et al., (2023). Curiously, the Review gives no attention to the serious limitations in the other reviews, even in instances where its own AMSTAR 2 ratings identified problems (e.g., Rew et al., 2021). This imbalance may suggest that reviews influencing restrictions in pediatric gender transitions received disproportionate scrutiny, raising further questions concerning analytic neutrality within the DRRC team.
Summary: The overview of the systematic review process lacks essential features of a rigorous systematic review: a comprehensive search, transparent and defensible selection criteria, structured synthesis of findings across reviews, and assessment of certainty. Instead, the analysis overlooks some of the highest quality English-language systematic reviews and largely limits its discussion to poorly justified criticism of the systematic review that supported policy restrictions in Sweden. Since the Utah Review authors did not provide an evidence synthesis, readers are left in the dark regarding whether the evidence base supports effectiveness, reveals uncertainty, or indicates potential harm.
6. Long-term outcomes
The Utah legislature requested an assessment of not only short-term outcomes of puberty blockers and cross-sex hormones in youth, but also of their long-term effects. When the University of Utah DRRC research team presented interim results of its evidence review to Utah’s DHHS in February 2024, it became apparent that the analysis lacked meaningful long-term outcomes, as most included studies followed patients for only 1–2 years after hormone initiation. The DRRC research team was therefore asked to conduct a dedicated analysis of long-term data, which is presented in Part II of the Utah Review.
For the “long-term” analysis, the University of Utah researchers focused on studies reporting outcomes of patients who, as a group, averaged at least 5 years post-treatment initiation. However, rather than conducting a new, comprehensive search designed to answer the question of long-term outcomes, they rescreened the previously identified studies focused on pediatric populations. The researchers identified 17 eligible studies that addressed the previously defined “high priority” outcomes and added a new “high priority” outcome: mortality.
The researchers did not conduct an analysis synthesizing and grading the certainty of the evidence. Instead, they supplied their narrative interpretation of benefits and harms as reported by the studies.
Mortality Risk
With respect to benefits, the researchers noted that some studies reported psychological improvements associated with hormonal treatment, including reductions in anxiety, depression, and social distress, as well as improvements in other measures of mental health and functioning, while other studies found no significant improvements.
With respect to harms, the researchers observed that some studies reported statistically elevated risks of cardiovascular mortality associated with ethinyl estradiol use, excess cases of certain benign brain tumors, and suicide and all-cause mortality outcomes that were higher than those of the general population, while other studies did not detect statistically significant increases in the harms assessed.
The “long-term” outcome analysis has numerous limitations that are summarized below:
- Inadequate search strategy. The researchers did not conduct a separate literature search to support the long-term outcomes analysis, asserting that the prior search was “comprehensive enough.” However, the original search was explicitly governed by pediatric search terms such as “Pediatrics,” “Child,” “Minors,” “Adolescent,” and “Puberty” (see Figure 2 below) and did not incorporate adult-specific search terms. As a result, studies reporting on adult cohorts were effectively excluded at the search stage. This created a structurally biased evidence base for long-term analysis, because most studies reporting long-term outcomes of hormonal treatments involve adult cohorts, and most of these studies never entered the potential study pool.
Figure 2

- Irregular study pool used for analysis. Despite structurally excluding adult studies at the search stage, some adult outcome data entered the long-term study pool, but only incidentally, when studies including pediatric patients also happened to report adult outcomes. Some included studies had virtually no pediatric patients and instead were comprised primarily of subjects whose treatments were initiated after age 18 (e.g., de Blok et al., 2021). As a result, the Utah researchers created an unbalanced study pool that allowed certain adult outcomes to inform conclusions about long-term outcomes while excluding others, based largely on happenstance rather than methodological principle. The researchers could have argued that outcomes for those treated exclusively as adults are not relevant to adolescents (a position with which we disagree but which is methodologically defensible) and excluded all adult treatment outcome data. Alternatively, they could have acknowledged that adult outcome data provide the most informative source of long-term evidence (with appropriate downgrading for indirectness). Instead, they selectively included some adult studies while excluding others, resulting in a narrow and uneven subset of adult data and a distorted evidence base.
- Lack of adherence to stated eligibility criteria. In addition to the problems with the eligibility framework described above, the screening process for eligibility appears to have been applied inconsistently, resulting in the exclusion of studies that met the stated eligibility criteria. For example, Klaver et al. (2020)—a 7-year follow-up study of adolescents—was not included, despite appearing to meet eligibility requirements. De Blok et al. (2019), was likewise excluded despite meeting the review’s criteria. Of note, De Blok reported a higher incidence of breast cancer among males treated with estrogen compared with the general male population. These exclusions raise concerns that study selection was not conducted in a systematic and consistently reproducible manner.
- Results dominated by a single patient cohort with questionable eligibility. Over 80% of the “long-term” sample comes from a single Amsterdam cohort (Wiepjes et al., 2020). That study includes patients at mixed stages of transition, including some who did not receive hormones. It also does not appear to meet the Review’s “long-term” threshold (≥5 years after treatment initiation), because it does not report the average duration of hormone exposure or specify how many participants had been on hormones for at least five years.
Inadequate assessment of risk-of-bias assessment due to confounding and co-interventions. Long-term hormone outcome studies are highly vulnerable to confounding: patients receiving cross-sex hormones often have poorer baseline mental health, socioeconomic disadvantage, and often receive co-interventions (e.g., psychotherapy or surgery), making it difficult to isolate the independent effects of hormones. These limitations do not appear to have been adequately accounted for in the Review's risk of bias assessments.
For example, the “Dutch Protocol” study (de Vries, 2014)—often cited as foundational because it is an early study reporting favorable outcomes—was rated as having a “fair” risk of bias for long-term hormonal outcomes. However, the study did not report outcomes following cross-sex hormone treatment alone, instead combining hormone treatment and surgery within a single follow-up assessment.However, the study text indicates that at least 3 of the 70 adolescents experienced adverse effects during the hormone phase (including severe diabetes and obesity), and that at least one apparent detransition occurred (de Vries, 2014). In addition, there were substantial missing data: at least 20% did not participate in the final phase, and over 40% lacked data on key follow-up mental health measures. These features suggest a critical risk of bias with respect to conclusions of safety and long-term psychological benefit.
- Absence of formal evidence synthesis. The DRRC team acknowledges that it performed no formal evidence synthesis and that its conclusions reflect the authors’ judgments after reviewing individual studies. Without structured synthesis or certainty grading, the analysis does not meet accepted standards for a systematic review and does not allow readers to assess the strength or certainty of the evidence regarding long-term benefits and harms.
- Failure to integrate long-term findings into overall conclusions. Evidence of harm emerging from long-term outcomes—particularly signals of elevated mortality—was not incorporated into the main conclusions presented in Part l. As a result, the Part I conclusion that “the evidence … supports that the treatments are safe in terms of changes to bone density, cardiovascular risk factors, metabolic changes, and cancer” (p. 90) remains unqualified by the longer-term mortality and morbidity findings presented in Part ll.
Summary: In addition to the absence of a structured evidence synthesis that assesses the certainty of the evidence, the Part II analysis of long-term outcomes suffers from additional serious limitations. One has to do with the exclusion of the most relevant studies of long-term outcomes of hormonally-treated adults. Since adults have been treated with cross-sex hormones for decades longer than minors, long-term outcome data are correspondingly more extensive in adult populations. Furthermore, adult treatment experience was also used to justify off-label extension of these interventions to adolescents, without robust additional research on pediatric-specific risks and benefits. A search strategy that systematically excludes adult long-term adult outcomes therefore cannot adequately support an assessment of long-term effects. As a result of this restricted search approach, several influential long-term cohort studies (e.g., Bränström and Pachankis (2019), Dhejne et al. (2011), and Jackson et al. (2023)) did not enter the analyzed study pool. Of note, these studies report elevated morbidity and mortality risks among hormonally treated adults and do not demonstrate clear long-term mental health benefit.
Moreover, the signals of harm that were identified in Part II itself were not incorporated into the Part I conclusions. Mortality, in particular, was analyzed only in the long-term section. Despite identifying elevated mortality and morbidity signals, the researchers did not revise or qualify their earlier safety conclusions, which continue to inappropriately state that no significant risks were identified.
7. Summary of Clinical Practice Guidelines
Although the requirement for “systematic medical evidence review” requested by the Utah legislature did not ask for an analysis of treatment guidelines for youth with gender dysphoria, and despite claiming to lack the resources to conduct an evidence synthesis (which the legislature did request), the University of Utah DRRC research team opted to provide an analysis of treatment guidelines.
The researchers used a broad definition of what constitutes a “clinical practice guideline,” yet found only five and assessed only four. Two are well-known clinical practice guidelines by WPATH and the Endocrine Society, and the other two are lesser-known guidance documents, including a relatively obscure ACOG (American College of Obstetricians and Gynecologists) position statement focused narrowly on menstrual suppression for gender-dysphoric females. This is an unusually small number; for context, other recent systematic reviews assessed 12 and 23 guidance documents on the same topic.
Of note, several peer-reviewed guideline quality appraisals that used standard tools found the guidelines summarized by the University of Utah researchers to be of poor quality and highly interdependent and circular. The most recently published appraisal of WPATH guideline quality, which instituted the strictest conflict-of-interest management protocols to date, found that recommendations relating to adolescents have serious limitations in “scientific and methodological rigor, applicability, and transparency in managing competing interests.” The researchers concluded that “uncritical adoption or endorsement of WPATH’s guidelines may result in a disservice or even harm to this vulnerable population.”
Included Guidelines
Applying the Johnston et al. (2019) framework to Utah’s guideline analysis shows it failed 17 of 20 criteria across all eight domains. No domain was rated adequate, including those covering quality assessment, data synthesis and analysis, and reporting and interpretation. Below, we provide a high-level account of the key problems that render the Utah Review analysis of guidelines unusable for evidence-based decision-making. The full appraisal using the Johnston et al. checklist is available here.
- Shifting guideline eligibility criteria. At first, the Utah Review appears to adopt the Institute of Medicine's (IOM) definition of a clinical practice guideline (CPG), requiring recommendations to be based on a systematic review of evidence and an assessment of benefits and harms (p. 13). Later, however, the Utah Review modified the IOM standard, stating that guidelines need only “cite and discuss” supporting literature so long as a systematic review was “performed” elsewhere in the document and the guidelines originate from a "recognized medical authority" (pp. 17–18). The motivation to relax the definition toward not requiring a systematic review for the recommendations is understandable. Without this requirement, neither WPATH's Standards of Care 8 nor the Endocrine Society guidelines would be eligible for inclusion in the analysis, as their recommendations for puberty blockers and cross-sex hormones in adolescents under 18 years of age do not cite any systematic reviews in support of their recommendations for this population. However, it also means that many other clinical guidance documents would also qualify. This makes it all the more surprising that the Utah review identified only five eligible guidelines—as compared to 12 to 23 guidelines identified and assessed by recent peer-reviewed analyses of guidelines on the same topic.
- Inconsistent adherence to eligibility criteria. Inexplicably, among the total of only five "qualifying" guidelines, two clearly do not qualify according to the University of Utah's modified IOM standard. Specifically, neither the European Society for Sexual Medicine (ESSM) Position Statement, nor the German 2013 guideline for child and adolescents meet the requirement of having performed a systematic review for any part of the clinical guidance document. Further, if the University of Utah researchers considered such consensus-based documents eligible, then additional clinical guidance documents—such as the influential American Academy of Pediatric's 2018 Position Statement and the Australian Standards of Care (Telfer et al., 2018)—issued by "recognized medical authorities" but entirely lacking any systematic reviews, would also have qualified, yet they are not included in the Utah analysis.
- Failure to consider qualifying international guidelines. While the 2013 German guideline was considered eligible by the Utah researchers—despite not including a systematic review of evidence—the more recent Swedish and Finnish national guidelines, which were underpinned by systematic reviews, were not included. Given the Utah legislature's specific request to analyze "literature from other countries"—and in light of Europe's leading role in reconsidering the practice of pediatric transition as an experimental interventions—it would be reasonable to expect for a proper guideline analysis to discuss—or at least identify— such important clinical guidance documents. Notably, Swedish and Finnish guidelines were the only two out of 23 recommended for implementation by the University of York systematic reviews.
- Cass Review. The 2024 Cass Review, together with its supporting systematic reviews, was published five months before the Evidence Review was submitted to Utah DHHS and more than a year before its eventual publication. The Cass Review is a major and influential report providing a foundational analysis of pediatric gender medicine. Several of its peer-reviewed systematic reviews are directly relevant, including those addressing puberty blockers, cross-sex hormones, and clinical guidelines, and it informed the adoption of more cautious treatment policies for gender dysphoria in the UK for children and adolescents. Despite this clear relevance and timing, none of these publications is mentioned in the 1,051-page Evidence Review. The Recommendation Report’s assertion that the DRRC “included an analysis of the Cass Review” (p. 4) is materially misleading, as it refers instead to presentations by Leor Sapir and Equality Utah (starting at 1:30:19 and 1:55:05 respectively) delivered to a Utah DHHS interim committee—work neither authored nor presented by the DRRC.
- No risk-of-bias or guideline quality assessment. Rather than using a validated instrument such as AGREE II to assess the methodological quality of clinical practice guidelines, the authors relied on the issuing body’s status as a “recognized medical authority,” substituting an appeal to authority for formal appraisal (p. 19). No established tool for assessing clinical guideline quality treats institutional prestige as a validated proxy for methodological rigor. This reliance is antithetical to the principles of evidence-based medicine, and is at the very root of the current problems in the field of pediatric gender medicine.
- Error in conclusion about the strength of guidelines. A central objective for Utah Evidence Review's analysis of guidelines was to report on the "level of evidence" (LOEs) that support guideline recommendations (p. 3). Not only was this objective unmet, but the researchers' conclusions are at odds with reality. Of the 4 guidance documents for adolescents they identified, 3 (including WPATH's Standard's of Care 8) were not supported by any formal evidence appraisal (rendering them non-evidence-based), while and the remaining one— by the Endocrine Society (ES)—was supported by evidence GRADEd as "very low" and "low" certainty. The University of Utah researchers failed to acknowledge the latter fact in their analysis, instead stating that the ES guideline as supported by all 4 levels of evidence—including "moderate" and "high" (p. 298). For the avoidance of doubt, the ES guideline never claimed that its recommendations for adolescents are supported by either "moderate" or "high" levels of evidence (ES guidelines, p. 3871). This appears more than an accidental oversight, as in their conclusions, the Utah reviewers incorrectly state, "high-quality guidelines are available to guide qualified providers in treating pediatric patients who meet diagnostic criteria." This level of misrepresentation of the evidence behind current treatment guidelines is alarming.
Following a non-systematic search which identified the total of four eligible English-language guidelines, the University of Utah researchers offered a descriptive summary of the recommendations, with a heavy emphasis on those issued by WPATH. This approach introduced multiple critical problems, including:
- Framing of puberty blockers and cross-sex hormones as standard of care. The researchers present treatment with puberty blockers and cross-sex hormones as the undisputed standard of care for gender-dysphoric youth. In addition to inappropriately framing these interventions as evidence-based, the Utah researchers frame them as “medically necessary” (p. 36).
- Emphasizing WPATH’s allowance for cross-sex hormones at the earliest signs of puberty. WPATH removed minimum ages for all medical and most surgical interventions in minors and permits cross-sex hormones at the earliest pubertal signs—potentially as young as 8–9. The Utah Review highlights this point, which remains under-appreciated.
- Suggesting that hormonal interventions are appropriate for adolescents with eunuch identities. Given the scope of WPATH SOC-8, the Utah researchers said they would focus only on recommendations relevant to minors (p. 13). In that context, their decision to discuss guidance on “eunuch” identity as eligible for hormonal therapy (p. 37) implies they viewed endocrine treatment for eunuch-identified pediatric patients as within the scope of what they characterize as standard of care, “medically necessary” interventions.
Making redundant claims that prepubertal minors do not receive hormonal interventions. The central controversies in pediatric gender medicine do not concern hormonal treatment of prepubertal children, who do not yet produce sex hormones in meaningful quantities. The Utah Review’s emphasis that prepubertal children do not receive hormonal interventions suggests an attempt to frame the “gender-affirming” treatment model as judicious and discriminating. Gender dysphoric minors become candidates for hormonal interventions beginning at Tanner stage 2, which can be as young as 8 or 9 years of age.
Summary: Synthesizing multiple clinical practice guidelines poses recognized methodological challenges: recommendations are often numerous and qualitative, terminology and scope vary, interventions may not align, and evidence-grading systems often differ—so identical labels can reflect different standards. As noted in Johnston et al. (2019), these complexities require explicit, systematic methods for comparison, synthesis, and transparent reporting. The Review does not overcome or even address these challenges. No framework for reconciling divergent grading systems is described, no structured meta-synthesis of recommendations is undertaken, and screening, selection, and extraction procedures are not comprehensively reported.
Instead of structured critical evaluation, the Utah Review offers what appears to be a set of cherry-picked guidelines identified through inconsistent application of shifting eligibility criteria, and accompanied by a selective descriptive summary of recommendations accompanied by interpretive commentary. Taken together, these features raise serious questions about methodological rigor, consistency, and neutrality in the guideline analysis, rendering the document unfit for evidence-based decision-making.
8. Desistance and Regret
The Utah legislature explicitly requested that the commissioned systematic medical evidence review analyze “rates of desistance and time to desistance”. The University of Utah DRRC team acknowledged their understanding that “persistence, desistance, and regret” were “pointedly of interest to the legislature,” yet explicitly stated that they would not consider these as “high-priority outcomes.”
At the same time, the Evidence Review claims to have provided an analysis by ‘examin[ing] the evidence in the retrieved studies” (p. 83). The Recommendation Report further bolsters this claim, stating that the “systematic review” provides “a complete summary” of the 32 papers in this section (Recommendation Report, p. 7).
However, a close reading of the Evidence Review reveals that the analysis of desistance studies in no way resembles a “systematic review.” Rather, the entirety of the analysis consists of this single passage, along with a table of 32 studies with a brief description of the findings:
“We found 32 studies that addressed persistence, desistance, and/or regret. Findings from these studies that relate to rates of persistence and desistance are summarized in Table I.26” (p. 83).
Such "analysis" cannot support the Review’s conclusion, “we found (based on the N=32 studies that addressed it) that there is virtually no regret associated with receiving the treatments, even in the very small percentages of patients who ultimately discontinued them.” (p. 91)
The Review's analysis of detransition is wholly inadequate at every stage—from the search criteria to study appraisal to drawing conclusions about the rate of detransition and the certainty of the estimates. We highlight some of the most glaring problems below.
Inadequate study identification and eligibility strategy. The University of Utah DRRC team did not conduct a targeted search for detransition outcomes. Instead, it relied on studies carried forward from its other analyses, which included studies identified using a search strategy that (1) did not include terms such as “detransition,” “desistance,” or “regret,” and (2) required an adolescent cohort. These design choices precluded meaningful identification of detransition studies, both because of the lack of relevant search term and because detransition and regret are typically found in studies of adults rather than adolescents, due to the well-established transition “honeymoon” period (the average time to detransition ranges between 5–10 years).
Even when detransition-related studies might have been captured incidentally under broader terms (e.g., “hormones”), the research team applied a second filter at the eligibility determination stage that screened out most of the remaining relevant studies. This is because eligibility was determined by whether a study addressed a “high-priority” outcome, and since “desistance” was explicitly excluded from the high-priority outcomes list, detransition-focused studies were screened out by design.
Weak indirect data. The deeply problematic search strategy is self-evident in the list of 32 studies identified by the Utah researchers. The studies specifically designed to assess the detransition phenomenon, which has become much more widely reported in recent years, are missing from the list (e.g., Boyd et al., 2022; Hall et al., 2021; Littman, 2021; Roberts et al., 2022; Sansfaçon et al., 2023; Vandenbussche, 2022).
At the same time, the 32 included studies capture only a small subset of detransition, desistance, and regret-related studies where these outcomes are assessed as part of a much broader outcome assessment, and with less focus on and precision around the outcome of detransition.
Lack of appropriate risk of bias assessment. The Utah researchers did not conduct a risk-of-bias assessment for detransition and regret outcomes since they did not classify these outcomes as “high priority.” As a result, they interpreted the detransition/regret evidence without a systematic evaluation of study limitations. This is problematic in any context, but it is particularly consequential here because studies of detransition and regret are often constrained by major definitional problems and other methodological limitations, as we have detailed in our prior analyses.
For example, the researchers cite Wiepjes et al., (2018) to support their conclusion that adolescent transition is associated with “virtually no regret.” However, the study used an extremely narrow definition of regret. Under its methodology, individuals would be counted as experiencing regret only if they had undergone gonad removal (ovaries or testes) and subsequently restarted natal-sex hormone supplementation. Using this definition, many well-known detransitioners currently suing their providers—such as Chloe Cole, Fox Varian, or the UK’s Keira Bell —would not be classified as having experienced regret. This illustrates how this and other studies use restrictive ascertainment criteria that likely underestimate the true incidence of regret.
- Errors in study summary statistics. There are numerous errors in the data tables. For example, in describing the Wiepjes et al., (2018) study discussed above, the researchers report the following incorrect statistics about pubertal suppression (PS) and cross-sex hormones therapy (HT): “No cases of regret were observed among the 1,360 individuals who were first seen before the age of 18 years. 1.9% of adolescents who started PS (n=812) stopped PS and did not start HT.” However, the number of adolescents who commenced PS is 333 adolescents, not 812. Further, regret rates only pertained to those adolescents who were on cross-sex hormones and underwent gonadectomy, which represents 309, rather than 1,360 adolescents.
- No synthesis of the evidence. The Review provides no synthesis of what the overall body of evidence shows—either narratively or quantitatively (e.g., using GRADE). Instead, readers are directed to Table 1.26 (pp. 83-89), which lists the (often error-ridden) characteristics of 32 studies, as demonstrated above. The Review does not caution readers about the degree to which follow-up duration and attrition can bias estimates downward or otherwise distort conclusions—despite the centrality of these limitations to interpreting the data it provides in Table I.26. The presentation of the studies in the table is also difficult to follow. To facilitate evaluation of this section, we manually compiled a fully referenced list of the studies used to support the Utah Review’s detransition analysis, available at this link.
Summary. The Utah researchers’ conclusion that detransition and regret are rare is not trustworthy. The best available data currently suggests that the rate of detransition among those treated in adolescence and young adulthood is between 10% and 30% within about 5 year post-transition. The Review could have reported this if their analysis had included the full range of relevant studies.
The University of Utah’s decision to de-prioritize desistance along with the related phenomena of detransition, persistence, and regret is difficult to justify. First, it conflicts with the Review’s legislative mandate by failing to address a question policymakers specifically sought to answer. Second, detransition and regret data can serve as early signals of potential harms in an intervention intended to be lifelong; excluding these outcomes weakens clinical quality and undermines informed consent. Third, the decision is internally inconsistent: despite deeming detransition “not high priority,” the Review attempted an analysis anyway—producing a flawed assessment from start to finish.
Against that backdrop, the Utah Review’s conclusion that there is “virtually no regret associated with receiving the treatments, even in the very small percentages of patients who ultimately discontinued them” is irresponsible at best—and plausibly reflects a dereliction of duty.
9. Interruption of Normally Timed Puberty
The Utah Legislature explicitly requested that the evidence review assess the “short-term and long-term benefits and harms of interrupting the natural puberty and development processes of the child.” The context for this question is interruption of natural and normally timed puberty in otherwise healthy adolescents for the treatment of gender dysphoria, rather than treatment of precocious puberty or other endocrine disorders.
Puberty is a time-sensitive, hormone-driven developmental process affecting bone, brain, fertility, and metabolism. Disrupting it may produce cumulative or delayed effects that short-term observational studies miss. For this reason, a credible safety assessment must integrate clinical data, long-term epidemiological evidence, and evidence from basic science and physiology on endocrine mechanisms and developmental timing. Such an analysis was performed in chapter seven of the U.S. HHS Evidence Review.
No such analysis of interrupting normally timed puberty has been provided by the University of Utah DRRC team. The following lists highlights some of the problematic decisions that led to an inadequate analysis of possible harms of interrupting normal puberty.
- Exclusion of “fertility” from the “high-priority outcomes.” Despite the Utah Legislature expressly requesting an analysis of fertility, the Review excluded it from its “high-priority” outcomes. The rationale for the exclusion was paradoxical: as the Recommendation Report (p. 4) stated, fertility was excluded precisely because hormone-related fertility harms are already a “known risk.” Excluding “fertility” as an outcome of interest virtually guarantees that one of the best understood and most ethically problematic effects of pediatric medical transition cannot be accounted for by the Utah Review’s review of evidence of benefits and harms.
- Failure to consider sexual function as an outcome of interest. Early puberty suppression followed by cross-sex hormones may impair sexual function. The lack of attention to sexual health further limits the Utah Review’s analysis of the long-term harms of interrupting normal pubertal development.
- Lack of consideration of neurodevelopmental effects of interrupting of normal puberty. Puberty drives structural and hormonal brain maturation, including synaptic pruning and myelination. No attempt was made by the University of Utah researchers to assess what is known about the effects of suppressing this biological developmental process on the brain.
The Utah Review offers no examination of these or numerous other potential harms of interrupting normal puberty. Instead, its conclusions assert that endocrine interventions such as puberty blockers and cross-sex hormones are safe. It even asserts elsewhere that “conditions like GD…have no other effective interventions” (p. 63). This demonstrably false and highly consequential claim is offered without citation or analysis and is treated as a settled fact.
Summary. The Utah Evidence Review omits the legislature-mandated analysis of the harms of interrupting natural puberty. It relies on the unsupported assertion (p. 63) that gender dysphoria has “no other effective interventions,” avoids examining domains where harms are anticipated (e.g., infertility), and fails to assess potential harms across other areas of physical, cognitive, and psychosocial development.
10. Recommendation Report
Following completion of the Evidence Review to Utah’s DHHS in August 2024, the advisors to DHHS submitted their individual reports—a process which culminated in the 16-page Recommendation Report published 9 months later—in May 2025. However, the process by which those individual analyses were aggregated into the Recommendation Report—and whether and how the consensus of the individual advisors was through—is opaque. No process is given on how a lack of consensus within the advisor group was managed or resolved.
While the Recommendation Report acknowledged some limitations in the Evidence Review, it inaccurately described the Evidence Review and made a number of other non-evidence based assertions, including:
- Misrepresenting the Evidence Review as a “systematic review.”
- Repeating the inaccurate assertions that the body of evidence shows hormonal interventions to be beneficial.
- Falsely equating not providing hormonal interventions to gender-dysphoric minors to not treating gender dysphoria and indicated without good evidence that withholding hormonal interventions may lead to “psychological and social harms.”
- Laying out recommendations for structuring the delivery of a hormone-treatment service for minors—without providing any suggestions for how to structure the care pathways in the event that the moratorium is not lifted.
The Recommendation Report makes four recommendations to the Legislature. In brief, the recommendations are:
- Create a hormonal transgender treatment board managed by DHHS which, among other duties, oversees appropriate provider training.
- Limit providers who can deliver hormonal interventions to demonstrated experts.
- Ensure that hormonal interventions are provided only within a comprehensive interdisciplinary care team model.
- Institute an enhanced and explicit informed consent and assent process.
We have identified significant shortcomings in the Recommendation Report that undermine the reliability of these recommendations:
- Recommendations not grounded in the Review’s own evidence base. The recommendations appear to have been formulated with little, if any, reliance on the Review’s own evidentiary corpus, instead turning to sources outside the document when evidentiary support was needed. Instead, the Recommendation Report cites publications not identified in the Evidence Review (two systematic reviews and a 2015 narrative review), as well as a scoping review that was incorrectly identified as an “eligible ‘systematic’ review,” but was ultimately not analyzed. Apart from demonstrating that the Review was not fit for purpose as a coherent decision-making document, these discrepancies indicate that the recommendations themselves rest on an opaque process for which no transparent evidentiary record has been provided.
- No explicit statement of values and assumptions. SB 16 required the department to describe the assumptions and value determinations underlying its recommendations, including whose values were considered, which benefits and harms were deemed critical, how they were weighed, how preferences may vary across patients and families, and how those judgments shaped the recommendation’s direction and strength. The Report does not provide this analysis. Instead, the DHHS lists broad principles—harm minimization, shared decision-making under uncertainty, and interdisciplinary, patient- and family-centered care—without identifying which values guided the recommendations, ranking patient-important outcomes, summarizing evidence on patient or family preferences, or linking any value judgments to recommendation strength.
- No specification of conditions under which treatment should not proceed. Although SB 16 expressly requires the department to specify when a treatment should not be permitted, the Report responds only with system-level conditions and does not define clinical thresholds, contraindications, or other circumstances in which hormonal treatment ought not to proceed.
- Inadequate guidance on informed consent content. The statute directs the department to recommend what information minors and parents should understand before consenting to treatment. The Report emphasizes enhancing consent and assent processes but does not identify the specific patient-important outcomes that must be discussed, distinguish established effects from areas of uncertainty, or connect required disclosures to the risk/benefit judgments underlying its recommendations.
Summary: The Recommendation Report summarized the Evidence Review’s findings and issued recommendations related to the provision of care to gender-dysphoric minors. Although the Report recognized some methodological limitations of the Evidence Review, it continued to incorrectly reference it as a “systematic review” and it accepted the Review’s non-evidence-based conclusions that puberty blockers and cross-sex hormones are safe and effective treatments for gender-dysphoric youth.
11. Interpretive Bias and Spin
Throughout the Evidence Review and the Report, there is evidence of interpretive bias, also known as "spin"—a phenomenon in research defined as "specific reporting that could distort the interpretation of results and mislead readers." Motivations for spin may vary, from unconscious confirmation bias to a more conscious desire to produce results that further a specific outcome.
There are multiple instances of spin in the Utah Review, a few of which are highlighted below.
- Biased introduction of the purpose of the Utah Review. The Introduction, “Gender Dysphoria and the Utah Context,” does not adequately present the clinical uncertainties in gender dysphoria treatments in youth as the basis for the Review, and instead focuses on political framing. Transgender identification is presented as a timeless phenomenon with a reference to the 1st Century BCE Roman collections of stories “about a transgender figure, Tiresias” (p. 2). Notably, the researchers omit the more recent and clinically salient exponential rise in youth identification. Further, while Utah Review acknowledges “public discourse” regarding whether adolescents should receive endocrine interventions, it frames the debate primarily through a political lens and does not mention the marked shift away from hormonal treatment of adolescents within the clinical community in a growing number of European countries.
- Proactive framing off-label use of "gender-affirming" interventions in youth as inherently safe. The researchers go beyond the requested analysis of the FDA status of hormonal interventions used to treat pediatric gender dysphoria, proactively framing the practice as inherently ordinary and requiring no added scrutiny. The researchers make flawed comparisons between the practice of using antibiotics not tested in pediatric populations to treat pediatric infections (which are short-term, administered to treat a well-defined disease, and proven in adult populations) and "gender-affirming" endocrine interventions (meant for long-term or lifelong use, administered in the absence of a well-defined disease, and never proven even for adults). The researchers' attempts to portray off-label prescribing as inherently safe by referencing to “one study,” while ignoring broader evidence linking off-label prescribing to higher adverse-event risk (risk/odds ratios ~1.67-2.25)—constitutes a prime example of "spin."
- Falsely presenting uncontrolled studies as though they had a control group. The Utah Review repeatedly treats uncontrolled studies as if they include valid comparators. For example, when outcomes are reported separately for males and females, the analysis states that one subgroup serves as a “control” for the other, even when the studies themselves made no such claims—such as the NIH-funded Chen et al., 2023 study which explicitly states: “our study lacked a comparison group." The Utah researchers, however, listed this study as having the "exposure" group as "AFAB"/females, and the "comparator" as "AMAB"/males. (p. 636). This manufactures a comparator where none exists and artificially inflates the perceived methodological strength of the evidence.
- Obscuring lack of improvement by presenting it as evidence of "stability." In discussing their results, the researchers write that “over time, treated transgender men showed reduced anger and anxiety, whereas treated transgender women were more stable” (p. 911). This phrasing functions as a linguistic sleight of hand: it presents improvement for one group in concrete terms (“reduced anger and anxiety”) while recasting an apparent absence of improvement for the other group as a positive-sounding outcome (“more stable”). The findings should be described transparently and symmetrically, without euphemistic language that can blur whether any meaningful change occurred.
- Misrepresenting quantity as quality. The researchers state that because their “highly exhaustive searches… yielded many studies that were not cited by many guidelines and SRs [systematic reviews],” the Utah Review “has the potential to more reliably capture the true consensus of the evidence compared to some of the other top-of-the-pyramid evidence summaries found in guidelines and SRMAs” (p. 91). This statement is remarkable, given the researchers’ own admission that they did not conduct a formal evidence synthesis or appraise the evidence for certainty. Either it reflects a profound misunderstanding of evidence-based medicine—which prioritizes the quality and certainty of evidence over the quantity of studies—or it functions to mislead readers into believing the Utah Review is more reliable than high-quality systematic reviews and guidelines that did apply evidence-based methods in their appraisals.
- Appeals to authority. The researchers treat institutional authority throughout the Review, often using it as a proxy for methodological quality. For example, in their analysis of clinical guidance documents from entities such as WPATH and the Endocrine Society, they state that they did not analyze methodological rigor because they “restricted inclusion to recognized medical authorities who published evidence-based guidelines [sic]” (p. 19). Paradoxically, as noted above, the Utah Review's analysis of clinical guidance documents, did not include the 2018 American Academy of Pediatrics (AAP) position statement despite its eminence—potentially because it was missing any references to a systematic review of evidence (per the Utah Review's criteria). Elsewhere in the Review, however, the authors cite support for endocrine interventions in minors by uncritically citing the 2018 American Academy of Pediatrics (AAP) position statement as “advocat[ing] for the importance of gender-affirming medical care in transgender adolescents, citing evidence that gender-affirming care may reduce depression, anxiety, eating disorders, self-harm, and suicide” (p. 2). This contradiction of whether the AAP statement represents credible guidance is never reconciled.
- Inappropriate leap from analysis to advocacy and recommendations. The practice of evidence-based medicine insists that evidence alone is not sufficient to make recommendations—including whether certain medical intervention should be offered widely, restricted, made available in research settings, or removed from offered treatments entirely. An explicit process of evaluating other factors—driven by the consideration risk-benefit ratio, quality of the evidence, and individual and societal values and preferences, and informed by additional factors such as resources, cost-effectiveness, feasibility, acceptability, and equity. The Utah Review authors do not engage in any of these additional analyses, yet state that "policies to prevent access to and use of GAHT for treatment of GD in pediatric patients cannot be justified based on the quantity or quality of medical science findings or concerns about potential regret in the future" (p. 91). This appears to be a significant overreach.
Reasons for spin may vary, ranging from ignorance about the subject to a conscious desire to shape a preferred narrative. Regardless of motivation, spin uses multiple tactics to lead readers to conclusions that exceed what the study design and data justify. Frameworks such as GRADE were developed in part to reduce this interpretive slippage by requiring conclusions to remain proportionate to the certainty of the evidence. By not following GRADE or a comparable framework, the Utah Review—commissioned and executed in a highly politicized context—made itself uniquely vulnerable to spin.
Although it was not within the scope of this analysis to ascertain the motivations for interpretive bias and spin throughout the Utah Review, in the process of reviewing the literature we came across a paper published by one of the advisors to Utah’s DHHS (Keeshin, 2024), before the Utah Review was finalized. The paper suggests the Utah Review may have been initiated with a specific policy goal: to "help maintain or open" access to hormonal interventions for minors.
The publication discusses the Utah moratorium on initiating gender transitions in minors as part of a broader trend across U.S. states and explains that—while the legislature did pass a bill restricting hormone interventions for adolescents with gender dysphoria—“the final version was significantly modified through a coordinated effort across advocacy organizations working with the bill sponsors.” The publication positions the Utah Review as a potential way to reopen the hormonal treatment pathway, stating that “as states pass adolescent bans on gender-affirming care across the country, Utah offers a potential pathway forward in restrictive states to help maintain or open access to care.”
Although initiating an evidence review with a policy goal in mind is not unprecedented in attempts to regulate the field of pediatric gender medicine (the HHS Review was also initiated following an explicit Executive Order by the U.S. Administration seeking to restrict minor transitions), it should heighten the scrutiny of the results to ensure that the researchers involved in the evidence review maintained their independence, conducted an unbiased analysis of the data, and reached conclusions grounded in the evidence. Unfortunately, the Utah Review fails to withstand such scrutiny.
12. Conclusion
The Utah Review, produced by the University of Utah DRRC research team, represents one of the most methodologically flawed assessments of evidence in pediatric gender medicine to date. The central requirement to deliver a trustworthy systematic review of evidence failed at a critical stage, as the researchers did not conduct an evidence synthesis that appraised the certainty of the evidence. Consequently, the Utah Review’s conclusions of significant benefits of pediatric transition and an absence of harms are not trustworthy.
While the Utah Review is incapable of supporting evidence-based decision-making in pediatric gender medicine, decision-makers currently have access to a number of trustworthy systematic evidence reviews of pediatric gender medicine, including several recent systematic reviews of puberty blockers and cross-sex hormones.
Some have mistakenly interpreted the Utah Evidence Review’s 1,051-page length as evidence of its rigor. While quantity is never a reliable indicator of quality, in the case of the Utah Review, the extreme length of the document is actually symptomatic of its methodological failings. Over 900 of the Utah Review’s pages would typically be treated as supplementary, e.g. tables of extracted study details and a compiled bibliography of all identified studies, (including those the authors deemed irrelevant and never analyzed). This reflects the researchers' own stated—but difficult to justify—claim that producing “a complete bibliography of all included studies identified in our systematic search… grouped into relevant publication types” was of “greater urgency” than completing a rigorous evidence synthesis (p. 4).
Even as a supplement, this material does not achieve its aims. It is dominated by poorly labeled tables and extensive redactions of study details impede meaningful engagement. The prominent placement of the redaction statement on the Review’s cover page—informing readers that the redactions were made “to protect the identities of pediatric transgender, nonbinary, and gender diverse patients”—suggests a performative intent. It is highly unusual to redact key study details in an evidence review and as a result of the redactions the document reads more like a heavily censored investigative file than a scientific report. The redactions are especially puzzling because those same study details are already in the public domain as published research. Moreover, published studies do not disclose identifying patient information; results are reported in aggregate, and individual cases are presented only with sufficient de-identification.
However, the Utah Review’s greatest problems lie not in its 900+ pages of virtually unusable material, but in the few dozen pages that are substantive. The analysis reveals a methodological process that failed from start to finish. After developing an unfocused study search strategy, the authors found themselves unable to properly assess the high number of heterogenous studies for methodological quality. To narrow the list of studies they had to review, they made an unjustifiable decision to de-prioritize important outcomes such as fertility; did not properly analyze either desistance or long-term outcomes studies; failed to incorporate signals of harm from long-term research into the main analysis; and did not properly assess the prevailing treatment guidelines. In addition, the researchers entirely skipped the critical step of synthesizing and evaluating the certainty of the body of evidence.
Despite these profound problems (and many other omissions and methodologically questionable decisions) the research team nonetheless concluded that the “consensus of the evidence” indicates puberty blockers and cross-sex hormones, when used for pediatric gender dysphoria, are beneficial and have no known harms. This conclusion is not only unsupported by the evidence they analyzed, but also contradicted by every credible systematic review to date.
III. SEGM Take-away
The Utah Review raises a broader concern: the continued production of documents that present a veneer of methodological rigor and scientific respectability yet are not merely inadequate but actively misleading. The field of pediatric gender medicine must accept the results of nearly a dozen quality evidence reviews indicating that the evidence of benefit remains highly uncertain, and that the evidence of harm, particularly when biologically plausible harms such as infertility are treated as clinically relevant outcomes, is comparatively much more certain. Reasonable people may disagree about the appropriate policy and clinical implications, such as whether these interventions should be reserved for exceptional circumstances, restricted to properly designed research trials, or decommissioned from routine clinical practice altogether, but in 2026 there is no longer serious scientific disagreement over the state of the underlying evidence.
Societal resources should be directed toward the work that would help resolve the clinical controversy over the best approach to treating pediatric gender dysphoria. This is the work of rigorously analyzing existing data on patients treated over the last two decades; producing higher-quality primary evidence on specific interventions where it is feasible; debating the ethics and appropriate boundaries of research in this domain; strengthening outcome surveillance and long-term follow-up; and confronting current disagreements regarding the proper role of adolescent patient preferences and the principle of autonomy when providing risky, life-changing interventions to otherwise healthy children struggling psychologically with their developing, sexed bodies.
What cannot continue, however, is the ongoing expenditure of societal resources on misleading studies, shoddy evidence reviews, and compromised guidelines, an expenditure which seems to have gone into high gear with the recent production, along with the Utah Review, of the flawed German “consensus” guideline, and several other similarly misguided and methodologically weak documents. Gender dysphoric youth deserve access to high quality evidence-based care—not ideologically focused efforts to distort the evidence in order to shield a deeply troubled area of medicine from scrutiny or correction.