Why preclinical peptide results often fail to replicate
An animal result and a human result are not the same claim at different strengths. They are different claims, and the published concordance research shows they disagree often enough that the animal finding cannot be treated as a provisional version of the human one. The causes separate cleanly into three. The model may not represent the human condition it stands in for. The experiment may have been designed, analysed or reported in a way that inflates the apparent effect. And the material the experiment was run on may not have been what the paper says it was. The first two are extensively documented across biomedicine. The third is specific to synthetic peptides and is, for anyone reading this literature, the least discussed of the three.
What did the concordance research actually find?
Perel and colleagues conducted a systematic review comparing treatment effects between animal experiments and clinical trials, restricting themselves to interventions where the human evidence was unambiguous, then asking retrospectively what the animal models had shown. The interventions were corticosteroids for head injury, antifibrinolytics in haemorrhage, thrombolysis in acute ischaemic stroke, tirilazad in acute ischaemic stroke, antenatal corticosteroids for neonatal respiratory distress syndrome, and bisphosphonates in osteoporosis (BMJ, 2007;334(7586):197; PMID 17175568, DOI 10.1136/bmj.39048.407928.BE).
The results did not divide into agreement and disagreement so much as into several distinct failure modes. Corticosteroids for head injury showed a favourable pooled effect in animal models — an odds ratio of 0.58 (95% CI 0.41 to 0.83) for adverse functional outcome — and no corresponding effect in clinical trials. Tirilazad is the sharper case: in animal models it reduced infarct volume by 29% (95% CI 21% to 37%) and improved neurobehavioural scores by 48% (95% CI 29% to 67%), while in patients with ischaemic stroke it was associated with a worse outcome. Antifibrinolytics reduced bleeding in clinical trials while the animal data were inconclusive, which is the inverse error. The authors concluded that the discordance may be attributable either to bias within the animal studies or to failure of the models to represent the clinical condition adequately.
Note what the tirilazad case establishes. The animal effect was large, statistically tight and directionally confident. None of those properties predicted the human result, and none of them are properties a reader can inspect a study for and be reassured by.
Is the problem the model, or how the studies are run?
Both, and the second is measurable. Kilkenny and colleagues surveyed the quality of experimental design, statistical analysis and reporting in published animal research, extracting detailed information from 271 publications describing work on live rats, mice and non-human primates at publicly funded UK and US establishments (PLoS One, 2009;4(11):e7824; PMID 19956596, DOI 10.1371/journal.pone.0007824).
Only 59% of the studies stated the hypothesis or objective together with the number and characteristics of the animals used. Most did not report randomisation (87%) or blinding (86%) — the two design features that exist specifically to limit bias in animal selection and outcome assessment. Only 70% of publications using statistical methods both described those methods and presented results with a measure of error or variability.
Those figures describe reporting, and a paper that does not report randomisation has not necessarily failed to randomise. But an effect size drawn from a study that cannot be checked for allocation bias, blinding or its own statistical treatment is an effect size of unknown reliability, and that is the position a reader is in for the majority of this literature. The same group subsequently published the ARRIVE guidelines to specify a reporting minimum (PLoS Biol, 2010;8(6):e1000412; PMID 20613859, DOI 10.1371/journal.pbio.1000412); the guidelines exist because the survey found what it found.
The scale of the downstream consequence has been costed. Freedman, Cockburn and Simcoe analysed the economics of reproducibility in preclinical research and put the cumulative prevalence of irreproducible preclinical work above 50%, corresponding to roughly US$28 billion per year of United States preclinical spending that does not reproduce (PLoS Biol, 2015;13(6):e1002165; PMID 26057340, DOI 10.1371/journal.pbio.1002165). Their framing matters as much as the number: irreproducibility is treated as a systemic property of the research pipeline rather than a matter of individual bad studies.
Where does the material itself come into this?
This is the part specific to synthetic peptides, and it is stated plainly in the analytical literature rather than the reproducibility literature, which is probably why the two conversations rarely meet.
D'Hondt and colleagues, reviewing related impurities in peptide medicines, catalogued what solid-phase synthesis routinely leaves behind: deletion and insertion sequences, diastereomeric impurities from racemisation, side-chain protection adducts, oxidation products, dimers and oligomers, residual counterions such as trifluoroacetate, and in some cases contamination by unrelated peptides. Their observation about consequences is the operative one here — these impurities "possess the capability of greatly influencing initial functionality studies during early drug discovery phases, possibly resulting in erroneous conclusions" (J Pharm Biomed Anal, 2014;101:2–30; PMID 25044089, DOI 10.1016/j.jpba.2014.06.012).
An in-vitro or animal experiment run on incompletely characterised material has a confound that no amount of randomisation or blinding addresses, because the confound is in the independent variable. If two laboratories obtain different results from nominally the same peptide, the material is one of the candidate explanations, and it is a candidate that can only be excluded by batch-level analytical data.
That is the narrow thing a supplier can contribute, and it is worth being exact about its scope. Our published certificates record a purity determination for a named batch on a named date by a named third party: BPC-157 batch 2026-03, GHK-Cu batch 2026-03, Retatrutide batch 2026-03, MOTS-c batch 2026-03 and NAD+ batch 2026-03. A batch identifier in a methods section is what makes an experiment traceable to a characterised lot at all. It does not make a result reproducible; it removes one reason it might not be.
What follows for reading the peptide literature?
Mainly that the evidence tiers should be kept distinct rather than averaged. An in-vitro finding, a rodent finding and a randomised human trial are separate categories of claim, and the concordance research above gives no basis for treating the first two as weak versions of the third.
Across the compounds most commonly discussed as research peptides, the evidence base is extremely uneven. A few have large randomised human trials behind them. Many have only animal and in-vitro work. Some have neither in any systematic form, and for those the accurate statement is that no human trial data exists — not that the question is open, and not that results are pending. Where a literature is thin, the thinness is the finding, and describing it as anything else imports a confidence the record does not contain.
Veridian Research supplies these compounds strictly for in-vitro laboratory research. They are not drugs and are not approved for human or veterinary use.
The short version
Direct comparison of animal experiments against clinical trials found substantial discordance, including one case where a large, tight animal effect preceded a worse human outcome. A survey of 271 animal publications found most reported neither randomisation nor blinding. Irreproducible preclinical research has been costed above 50% prevalence and around US$28 billion a year in the United States alone. And in synthetic peptides specifically, uncharacterised material is documented as a source of erroneous conclusions in early functionality work. None of these is a reason to discount preclinical evidence. All of them are reasons to state which tier a finding belongs to.