SCIENCE

Berkeley Lab Catalyst Papers Split on How the Controls Were Run

Two papers from the same institution can report oxygen evolution activity for the same nickel-iron oxide that differs by a factor of five. The disagreement rarely comes from the material. It comes from how each group measured the active surface and corrected for the resistance between the working and reference electrodes. This piece walks through the procedural choices that produce the split and what a careful control actually requires.

The Reproducibility Split in Catalysis

Electrocatalysis has a benchmarking problem that materials science shares with fields as far apart as genomics and astronomy: the measurement pipeline is part of the result. When a group reports a current density at a given overpotential, that number depends on the geometric area of the electrode, the loading of catalyst, the electrolyte purity, and the way resistance is subtracted. Change any one and the ranking of materials can flip.

At Berkeley Lab and its affiliated user facilities, groups working on similar oxide catalysts have published activity metrics that are difficult to reconcile. One camp runs what might be called routine protocols: a quick scan, a standard reference electrode, a geometric area normalization. Another camp treats each measurement as a surface-science experiment, with electrochemical surface area (ECSA) determined separately for every sample and full iR correction reported in the main text.

The dispute centers on ECSA. If two labs disagree about how much active surface a catalyst has, they will disagree about its intrinsic activity even when the raw currents match. The same material, different reported activity, is the normal outcome.

Why ECSA Measurements Diverge

ECSA is the portion of a catalyst's surface that actually participates in redox reactions. The geometric area, the simple length times width of the electrode, does not correspond to that active area. For platinum and other noble metals, hydrogen underpotential deposition is the traditional probe: the charge associated with adsorbed hydrogen gives a surface count, provided the surface is clean and the crystal facets are known. Cleanliness is the hard part.

For oxides, the situation is worse. Double-layer capacitance is often used as a proxy, on the assumption that capacitance scales with active area. That assumption holds only if the double-layer behaves ideally, which it rarely does on a porous, defective oxide. Probe molecules such as adsorbed ions or redox species can poison active sites, so the act of measuring the surface changes it.

There is no single standard for oxide catalysts. A group can choose hydrogen underpotential deposition, double-layer capacitance at a fixed scan rate, or a redox titration, and each produces a different surface area. The error bars on site counting are large enough that turnover frequencies computed from them carry real uncertainty. A related piece on this site about how a hippocampal replay study shrank under larger samples makes the same point in a different field: measurement choices travel with the number.

The Role of Reference Electrodes

Reference electrode calibration is where small errors become large ones. A silver-silver chloride electrode that has drifted a few millivolts since its last calibration will shift every reported potential. A few millivolts is negligible at high overpotentials and decisive near the onset of a reaction, which is exactly where many comparisons are made.

Luggin capillary placement adds another variable. The capillary tip must sit close enough to the working electrode to minimize the uncompensated resistance, yet not so close that it shadows the surface. Move it a millimeter and the iR drop changes. Some labs report potentials against a reversible hydrogen electrode, converting on the fly. Others report against Ag/AgCl and leave the correction to the reader.

Neither choice is wrong, but they are not interchangeable. A paper that states potentials against Ag/AgCl without a calibration note forces the reader to guess. This site has argued in a piece on spectrograph pipelines that provenance decisions made early determine what later analysis can recover, and electrode calibration is the electrochemical version of that.

Surface Area or Intrinsic Activity

Geometric current density is the easiest number to report and the easiest to misread. If one electrode carries ten times the catalyst loading of another, its geometric current will be higher even if the material is less active per site. Comparing geometric current densities across papers is a common error, and it favors high-loading electrodes that may be mass-transport limited.

Turnover frequency is the more honest metric, but it requires counting active sites. Site counting via titration or capacitance has error bars that can span a factor of two. Reviewers often ask for both geometric and intrinsic metrics precisely because neither alone settles the comparison. Reporting both costs space and forces the authors to show their work.

There is a real trade-off here. Intrinsic activity is the scientifically meaningful quantity, but it is also the one most sensitive to the ECSA method. A group that reports only geometric activity may be accused of hiding behind loading, while a group that reports only turnover frequency may be accused of burying a questionable site count. The objection from the geometric camp is that intrinsic metrics are not comparable across labs when the site-counting method differs, and that objection is correct.

What a Careful Control Looks Like

A blank run with a bare electrode establishes the background current and the contribution of the substrate. Without it, a small catalytic signal can be indistinguishable from the electrode's own response. The blank should be run in the same cell, with the same electrolyte, on the same day.

Using the same batch of catalyst across all tests removes batch-to-batch variability from the comparison. If a paper compares three materials, each should be measured under identical loading, identical electrolyte, and identical scan parameters. Full iR-corrected polarization curves belong in the main text, not only in the supporting information, because the correction changes the shape of the curve near onset.

Electrolyte purity and pH should be stated explicitly. Trace iron in nominally pure potassium hydroxide can turn an inactive nickel electrode into an apparently good catalyst, a well-known effect in the field. A careful paper reports the electrolyte source, the purification procedure if any, and the measured pH.

Steps for Reading Catalysis Papers

The next time a catalysis paper crosses your desk, these checks take ten minutes and change how much weight the activity number deserves.

  1. Find the ECSA method in the supporting information. If it is double-layer capacitance, check whether the scan-rate range is wide enough for a linear fit and whether the capacitance is reported per geometric area.
  2. Look for reference electrode calibration details. A stated calibration against a reversible hydrogen electrode, or a fresh Ag/AgCl measurement, is a good sign. Silence is not.
  3. Compare geometric and intrinsic activity side by side. If only one is reported, ask why the other was omitted.
  4. Note whether blank runs, batch consistency, and electrolyte purity are described. Their absence does not invalidate the result, but it lowers confidence.
  5. Request raw data when a comparison hinges on a small difference. A code-archived result can still be rejected on provenance grounds, as this site described in a piece on a preprint verdict that turned on provenance, and the same standard applies to a polarization curve.