Article

Why methylene blue trials are difficult to blind

Why methylene blue trials are difficult to blind

Thirteen volunteers each swallow an opaque capsule, lie in a scanner, and perform memory tasks. Nobody in the room knows who received the drug. The trial report calls the design double blind, and at that moment the description is accurate. Then everyone goes home, and every participant who received methylene blue voids blue urine. Was the trial blind? The honest answer is narrower than the label: it was blind when some measurements were taken, and not blind when others were. That distinction is the entire subject of methylene blue clinical trial blinding.

Concealment at dosing is not masking through follow up

Blinding, which methodologists increasingly call masking, means that the people who could bias a result do not know the treatment allocation. Three groups matter: the participants, the staff who interact with them, and the people who assess the outcomes. Randomisation plus an opaque capsule conceals the allocation at the moment of dosing. That is where most readers stop thinking, and that is the mistake.

A blue dye leaks information over time, and each leak lands on a different group at a different moment. The capsule hides the colour at swallowing. Hours later the urine reveals it to the participant. Monitors can reveal it to staff sooner. So the correct question about any methylene blue placebo comparison is not whether the trial was blind in general. It is who knew what, and when, relative to each measurement.

What gives the allocation away

The tells are physical, dose dependent, and well documented:

Blue or blue green urine. This is the main tell at low oral doses. It appears within hours of ingestion and typically persists for about a day. In the San Antonio functional imaging trial, every subject who received methylene blue reported voiding blue urine after leaving the imaging centre, while no placebo subject reported any discolouration. The urine finding arrived after the scan, which matters, as will become clear below.

Blue stool and a blue tongue. These appear briefly after oral dosing and are visible to participants and to anyone examining them.

Skin and sclera tinting. At higher intravenous doses the dye visibly colours skin and the whites of the eyes. This tell reaches staff during the visit, not only the participant afterward.

Sensor interference. Methylene blue absorbs the light that pulse oximeters use, so oxygen saturation readings become unreliable during infusion. In the King's College cerebral blood flow study the team could not monitor oxygen saturation for exactly this reason, which means bedside staff had an instrument level signal that something was different.

Note the asymmetry: everything the dye reveals points in one direction. A participant with blue urine learns they received the drug. A participant with normal urine learns only that they probably received the methylene blue placebo, unless the placebo itself was designed to colour urine, which the ones tried so far do not.

What trials have actually tried

The methods sections of the imaging trials make the best case study, because they describe their countermeasures in unusual detail. Readers following the low dose evidence will recognise these studies from the memory literature.

The two San Antonio trials randomised healthy adults to 280 mg of oral USP grade methylene blue (about 4 mg/kg) or a placebo of FD&C Blue No. 2 food colouring, packed in opaque immediate release capsules. The coloured placebo matched the capsule contents at dosing but did not colour urine, so the mimic failed exactly where it counted. Knowing this, the Radiology trial team added a behavioural protocol: participants were asked to urinate before taking the capsule and not again until after the imaging session, specifically to avoid compromising the blinding through urine colouration. A separate research nurse held the randomisation key in sequentially labelled containers and took no other part in the study, and blinding was maintained for participants, nurses, and outcome assessors until the analysis was finished, according to the multimodal imaging report. The companion connectivity study used the same design and reported the same aftermath: only the methylene blue group noticed transient urine discolouration after leaving the centre.

The King's College study took a different and more candid route. Eight healthy volunteers each received a glucose placebo infusion and two intravenous methylene blue doses (0.5 and 1 mg/kg) on separate days in randomised order. The design is stated as single blind and within subject: investigators knew what was infused, and each participant served as their own comparison across sessions, as described in the blood flow and metabolism paper. With an intravenous blue dye there was no pretence that participants would stay masked across visits.

Older psychiatry literature shows a third strategy: the low dose as active control. A two year crossover trial of methylene blue for manic depressive psychosis used 15 mg per day as the control condition against roughly 300 mg per day, a detail recounted in the discussion of the blood flow and metabolism paper. Later work on a methylene blue derivative for Alzheimer disease used a low dose originally intended as the control and found it behaved like a treatment. A low dose control mimics side effects and visible tells far better than an inert capsule, but it concedes something important: the trial can only estimate the difference between doses, never the effect of the drug against nothing.

The sharpest test of the coloured placebo idea comes from a PTSD trial that pitted 260 mg of oral methylene blue against indigo carmine placebo in visually matching capsules. Participants were even encouraged to drink water to dilute the effect. Urine discolouration was still reported by 95.6 per cent of the methylene blue group against 21.1 per cent of the placebo group, in Telch and colleagues' extinction trial. A dye in the capsule matches the dosage form. It does not match the participant experience.

Trial method comparison: what each party could observe

This table is the article in compressed form. Read each row as a timeline, not a grade.

TrialDose and routeStated designPlaceboWhat participants could observeWhat assessors could observeWas masking success tested?
Rodriguez et al., Radiology 2016 (n=26)280 mg oral USP methylene blue, single doseRandomised, double blind, placebo controlledFD&C Blue No. 2 in opaque capsulesBlue urine after leaving the centre (all drug recipients); nothing during the scan, aided by the no urination protocolNothing during scanning; outcome analysis stayed blinded until completionNo guess test; urine reports collected as a delayed drug marker
Rodriguez et al., Brain Imaging Behav 2017 (n=28)280 mg oral USP methylene blue, single doseRandomised, double blind, placebo controlledFD&C Blue No. 2 in opaque capsulesTransient urine discolouration after leaving the centre, drug group onlyNothing during scanning; participants and investigators blinded until analysis concludedNo guess test reported
Singh et al., J Cereb Blood Flow Metab 2023 (n=8)0.5 and 1 mg/kg intravenous, plus glucose placebo, crossoverSingle blind, within subject, randomised order50 mL glucose solutionBlue urine and infusion related sensations across sessions; masking across visits implausibleInvestigators unblinded by design; imaging endpoints quantified computationallyNo guess test; subjective energy, mood, and pain ratings reported per session

Two patterns stand out. First, none of the three ran a formal check of masking success, such as asking participants to guess their allocation at exit. Stated blinding is a description of procedure; a guess test is evidence of outcome, and that evidence is absent. The gap is general, not specific to this dye: a review of 191 placebo controlled trials found that only about 2 per cent assessed blinding success in both participants and investigators or assessors, as reported by Fergusson and colleagues. Second, timing rescues part of the claim in the oral trials: the tell arrived after the scanner data were collected, so the imaging endpoints were measured while masking plausibly held, while anything rated later was not.

Testing guesses is possible, and one methylene blue trial shows how. In the intradiscal injection trial for discogenic pain, patients were prevented from seeing the injected solution and were asked afterward which treatment they thought they had received. About three quarters in each arm answered that they did not know, with no significant difference between groups, in the Kallewaard trial report. That is qualitatively different from writing double blind in a methods section. Formal indices exist for exactly this purpose, notably the Bang blinding index described by Bang and colleagues.

Why expectations and assessors matter unequally

Once a participant works out their allocation, expectation effects concentrate on subjective endpoints: energy and mood ratings, symptom scales, effort dependent tasks. The King's College subjective ratings illustrate the exposure. Participants rated energy, mood, and pain after each infusion, at a point when anyone receiving the dye had ample reason to know it. The paper reports higher energy ratings at the lower dose and pain ratings rising with dose. Those readings may be pharmacology, expectation, or both, and the design cannot separate the two.

Objective endpoints resist expectation differently. A participant cannot will a blood oxygenation pattern into a scanner, which is one reason the imaging results keep much of their value even with imperfect masking. But objectivity is not immunity. Analysis pipelines involve region selection, thresholds, and model choices, and an unblinded analyst can steer those choices toward a favoured result. The San Antonio teams closed that hole by keeping the blind intact through the analysis stage. That single sentence in their methods carries more weight than the word double blind in their abstracts.

The standard toolkit for dye like interventions follows directly: keep the people who assess outcomes separate from the people who administer the drug, use standardised scripts for participant contact, choose primary endpoints that are measured before the tell where possible, and run an exit guess questionnaire with a proper index rather than asserting that masking held. The memory trial results need exactly this kind of endpoint by endpoint reading.

How to read a blinding claim in five minutes

Apply this checklist to any methylene blue trial, and treat the answers as modifying each endpoint separately rather than passing or failing the whole paper:

  1. What exactly was the control: inert capsule, coloured placebo, or low dose? A coloured capsule that does not colour urine matches appearances only at dosing.
  2. Was the allocation concealed at dosing: opaque capsules, masked infusions, independent randomisation?
  3. Did the authors acknowledge the dye tells, and did they add any protocol against them, such as scheduled urination or separate staff roles?
  4. Who was stated blind, and until when: dosing, assessment, analysis? Blinding that ends before analysis is weaker than blinding that survives it.
  5. Was the primary endpoint subjective or objective, and was it measured before or after the tell? In scanner data and post visit ratings live on opposite sides of that line.
  6. Did anyone test whether masking succeeded: guess rates, a blinding index, anything beyond the design label?

The dye does not make fair trials impossible. It makes the single word blind insufficient as a summary. The way the dye behaves in the body guarantees that information leaks, and the safety profile literature confirms the tells are normal pharmacology rather than accidents. So read the methods section as a timeline of who could see what and when. Where the measurement came before the blue urine, trust it more. Where it came after, discount accordingly, and require replication with endpoints the dye cannot reach.