Article
Methylene blue reviews: what personal reports can and cannot tell you

Type “methylene blue reviews” into a search box and the reports arrive before the trials do. One person writes that their afternoons came back. Another writes that nothing happened. A third describes a bad night and stops posting. The tempting move is to count them: if enough people describe the same change, something must be there.
A collection of personal reports is not a small clinical trial. It is a different kind of object, assembled by people who never agreed to be a sample.
What a single report contains
Consider a realistic post: “Two weeks on methylene blue, ten drops in water each morning. Brain fog gone, I feel like myself again.”
Four claims sit inside those three sentences. A product was involved, an amount was used, a duration passed, and something changed. We read the last claim and treat the first three as decoration, when they are the limitations. The product is whatever the bottle held. The amount is a count of drops, which is not a measurement. The change is a subjective impression, rated by the person who expected something to happen, with no comparison and no baseline record. That is one observation, and it is real, since the person did feel different. It is not yet an explanation.
Four other explanations fit the same sentence just as well: expectation, something else changing at the same time, the symptom's own fluctuation, and the measurement drifting. None can be excluded from the report alone.
The people who felt nothing have less reason to write
The clearest evidence that online reports are a distorted sample comes from a 2014 comparison of 1,901 Amazon reviews of weight-loss and fertility treatments against trial outcomes. After six months on the Atkins diet, 93% of online reviewers (64 of 69) reported losing at least 10 kg, while 27% of trial participants (19 of 71) achieved the same change. The authors' conclusion is the sentence to remember: “Online reviews overestimate the benefits of medical treatments, probably because people with negative outcomes are less inclined to tell others about their experiences.” Someone who improved has a story to tell; someone who spent six weeks and felt exactly the same has nothing to report.
The direction of the distortion is not fixed, which is instructive. A 2025 analysis of 369 Reddit posts about fertility supplements found the opposite balance: among 279 authors who reported taking supplements, 9.3% described a positive perceived effect and 12.9% a negative one, while 21.1% asked for advice. For a forum built around troubleshooting, that is the expected pattern. A separate study found that patients who joined an online health community had higher depression and anxiety scores than non-members.
The mix of reports is a property of the community, not a census of outcomes.
Expectation is an ingredient
The usual objection is personal: “I am not the sort of person who is fooled by expectation.”
In 2010, a Harvard-led team gave 80 people with irritable bowel syndrome either placebo pills described to them honestly as inert, or no treatment at all, with equal practitioner contact in both arms. At three weeks, global improvement scores were 5.0 ± 1.5 for the open-label placebo group against 3.9 ± 1.3 for the untreated group. A later meta-analysis of eleven trials covering 654 participants found a pooled effect of 0.72 standard deviations. Someone taking methylene blue is not even in the honest-placebo position: they believe it is active, they chose it, they paid for it, and they are the one recording the result.
Physiological measurements move too. In a balanced-placebo study in runners, 1000-metre time trials improved only when participants believed they had taken caffeine, whether or not caffeine was in the capsule: effect size 0.43 for the deception condition against 0.42 for the real thing. Being told they had taken a placebo produced no improvement even with caffeine on board.
Nobody can hide the colour
Most supplements let a friend prepare identical capsules. Methylene blue does not: taken by mouth, it colours urine blue or green, and sometimes faeces.
A 2016 randomised trial in Radiology handled this carefully. Participants received either 280 mg of oral USP-grade methylene blue or FD&C Blue No. 2 food colourant in identical opaque capsules, and were asked to urinate before dosing but not afterwards until scanning ended, to protect the blinding from urine colour. The attempt is documented, and so is its failure: all participants who received methylene blue reported blue urine afterwards, and none of the placebo group did. The authors called the discolouration “a delayed marker confirming that they received the drug.”
A larger example shows the cost. The 2016 phase 3 trial of LMTM randomised 891 patients with Alzheimer's disease and had no placebo arm at all. Its control group received 4 mg twice daily, chosen, the authors state, “to maintain blinding with respect to urine or faecal discolouration.” An inert tablet would have been visible in the toilet. The bipolar-disorder crossover study that enrolled 37 patients and labelled its 15 mg arm “placebo” made the same compromise against an active 195 mg arm.
Read those designs as a warning about personal experiments. If hundreds of patients in a funded trial, with opaque capsules and a statistician's randomisation list, could still identify their assignment, then someone measuring drops into a glass of water at home already knows what they are taking. Blinding is not difficult here. It is unavailable.
The symptom has its own schedule
People rarely begin a supplement during a good week. They begin during a bad one, which introduces regression to the mean: extreme values drift back toward average ones whether or not anything is done about them.
The drift is measurable. In an untreated group of chronic low back pain patients, mean pain on a 0 to 100 scale went from 56 at the first visit to 54 two months later and 46 after four months. The authors attribute part of that movement to the attention of being studied, another confounder present in any self-experiment.
Spontaneous improvement is common in the conditions people cite when discussing methylene blue. Among 79 patients with chronic fatigue syndrome of less than two years' duration and no systematic intervention, 46% reported improvement at one year. In a cohort of 1,266 people hospitalised with COVID-19, about 65% of those reporting brain fog, memory loss, or concentration loss at a mean of 8.4 months no longer reported it at 13.2 months, while others developed the symptom in the same period. In untreated major depression drawn from primary care, an estimated 23% remit within three months and 53% within a year.
None of this means methylene blue does nothing. It defines the baseline a personal report does not have, and a change of the size usually described in a forum post falls inside the range that happens on its own.
The score has its own schedule
Repeated measurement changes the measurement. Healthy adults who completed a large neuropsychological battery seven times in a year improved with nothing but repetition: practice effects of 0.36 to 1.19 standard deviations during frequent testing up to month three, after which performance plateaued. The largest jump came at the second sitting. A rising score across three weeks of daily brain-training tests is the expected result of taking the test three weeks in a row.
Self-rated scales are not immune. Repeated Beck Depression Inventory administration in a student sample produced a 40% decline in scores over eight weeks attributable to repeated measurement alone, accounting for about 10% of the variance. “I have been rating my mood every morning and it is going up” is a sentence about the rating as much as about the mood.
Something else usually changed at the same time
Most people writing a methylene blue report were already changing other things: sleep, exercise, alcohol, a second supplement, daylight, a stressful project that ended. The report names the product because it is the newest item in the list, not the only one.
A report does not carry a verified amount
Drops are the standard unit in methylene blue reviews, and they are not a unit. A study of 19 multi-dose dropper bottles found single-drop volumes ranging from 33.8 to 63.4 microlitres, close to a two-fold difference between bottles, with shallower angles producing smaller drops. “Ten drops” from one bottle can be nearly twice “ten drops” from another, and the post usually omits the concentration, so even a well-measured volume cannot be converted into a mass. The guide to dosage units works through what that conversion requires.
Amount also matters in a non-obvious way, because the dose-response is not a straight line. In rat experiments, 4 mg/kg improved behavioural habituation and object recognition, 50 to 100 mg/kg reduced running-wheel activity, and low concentrations raised brain oxygen consumption in vitro and after dosing in vivo, a pattern the authors describe as dose-dependent in both directions. The memory and focus review separates human studies by dose for the same reason. A positive review therefore does not identify an amount that would work for you, and an account of feeling nothing does not rule out a response at a different amount.
There is a further gap between amount and exposure. In a crossover study of 16 healthy adults, mean absolute bioavailability of an aqueous oral dose was 72.3% with a standard deviation of 23.9 percentage points. Two people taking the same measured amount need not reach the same blood concentration, as the absorption and distribution review sets out.
A report does not carry a verified product
Every review is a review of a bottle, and measured content often diverges from the label. An analysis of 30 melatonin products found more than 71% outside 10% of the stated content, from 83% below to 478% above the claim, with lot-to-lot variation reaching 465% and serotonin in eight of them. In a collection of 272 herbal and dietary supplements implicated in liver injury, 51% were mislabelled, rising to 82% among steroidal products.
Methylene blue has a specification of its own, and material made for dyeing is not held to it. A product conforming to the pharmacopoeial monograph is tested for identity, purity, organic impurities such as Azure B, residual solvents, elemental impurities, residue on ignition, microbial limits, and bacterial endotoxins. The impurities and test-methods article shows those tests and what published sample results for methylene blue do and do not cover.
This is the part of the buying decision a reader can verify. Many sellers of methylene blue that advertise conformance to USP test only for heavy metals and describe that as full compliance. Blupreme works with a pharmaceutical manufacturer of methylene blue and publishes batch-specific certificates of analysis linking the batch number on the bottle to the tested ingredient batch. The USP designation page explains what the specification covers, and the guide to reading a certificate of analysis shows how to match a report to the lot you received.
Knowing what was in the bottle removes one confounder. It does not turn an unblinded single-person report into evidence that the product caused the change.
What reports are still worth reading
Personal accounts are poor evidence of benefit and reasonable evidence of something else: a possible problem. If several people independently describe the same distinctive reaction, that is a hypothesis for a clinician, and the side effects overview explains how to describe an exposure and where a suspected reaction can be reported. The overdose and poisoning guide covers the urgent end of that.
Reports also tell you which questions to ask. A thread of people disagreeing about whether methylene blue helped their sleep is a signal that sleep is worth looking up in the trial literature, not a signal that methylene blue helps sleep. The human evidence review shows what that literature contains, and the research library collects the trials.
Reading a thread without fooling yourself
The routine is short. Copy the claim exactly. Separate what the person observed from what they concluded. Then list the candidate explanations and ask which one this report rules out. In most cases the honest answer is none, and the report's status is unresolved rather than true or false.
Download the anecdote evaluation worksheet to run that check on a specific thread. It lists the fields that are usually missing, a place to record the rival explanations, and a log format for anyone who decides to try something anyway.
If you do test something on yourself, four choices make the record more useful than the average post: fix the outcome and the rating scale before the first dose, keep a baseline long enough to show how the symptom normally moves, change one thing at a time, and record the days that went badly. The guide to checking a protocol found online applies the same discipline to instructions rather than experiences.
Two symmetries close the argument. Feeling better does not establish that a product caused the improvement, and feeling nothing does not establish that the product does nothing at any amount or in any form. Both errors come from treating a single unblinded observation as a result. What reviews can add is a reason to look, and a warning worth acting on.