• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Fake science is out of control

From the Academy's by-laws:

The only criteria for membership shall be the quality and extent of their already published publichealth work. There shall be no discrimination based on gender, age, marital status, sexual orientation,religion, disability, ethnicity, skin color, nationality, place of residence or academic affiliation.​

In other words, in contrast to other scientific honor societies (eg, NAS, AAAS), that now give preference to women and minorities in their selection for membership, this new academy will select members strictly on their scientific merits.

Bravo.
This thread is about fake science, not fake egalitarianism. But hey, never let an opportunity go by, eh?
 
I would assume that AI is singularity bad at this.
For one, it would be easy to flood the training data with with bad examples (it as is the case now).
Secondly, science today is so pigeonholed that one would have to chose the training set extremely narrowly.
Third, fake science is mostly used to get grants and promotions, not to prove the existence of of the supernatural: the claims are usually not outrageous unless you know the subject very well.
And lastly, different countries, laboratories, topics of research and methodologies have vastly different approaches to the number of co-authors on a publication: just because one paper might have 2 and another 20 doesn't meant that one is more likely to be made up than the other.

Much better would be to pay scientists for successful debunks.
It would be trained on a dataset that consisted of good papers and faked/withdrawn papers and it would be "told" which were which.

I do wonder if Google's new "co-scientist" AI could already be used to detect fake/withdrawn/bad papers.
 
One of those things where the effect is likely to be opposite to the stated goal. These are people whose actual jobs are to identify "fraudulent and wasteful spending" so that they can "put an end" to it.
 
Wonder if this is a good area for current AI approaches to deal with? Spotting patterns and similarities in large data is something they are very good at.
Already has by some of the debunking groups. I think possibly Data Collada as well. And you are right! ai is great at the mind numbing, watching paint dry, task of data categorization and analysis
 
One of those things where the effect is likely to be opposite to the stated goal. These are people whose actual jobs are to identify "fraudulent and wasteful spending" so that they can "put an end" to it.
You can hear some of those people here:
They Were the Original DOGE. Then Trump Fired Them. (New York Times on YouTube, Mar 8, 2025 - 10:30 min.)
President Trump has sworn to root out corruption within the government, yet one of his first acts as president was to fire over a dozen independent watchdogs who did exactly that. We spoke to seven of them about the abuses they uncovered, what they really think about DOGE and what all this means for the future of American democracy.
0:00 — Intro
2:35 — What’s an inspector general?
3:20 — You got fired. Why should we care?
4:59 — What worries you about the future?
6:31 — What’s do you think about DOGE?
8:41 — Is our democracy in danger?
DOGE appears to enable corruption and tax fraud rather than prevent it.
 
Just an example of the kind of stuff that could apparently be published in academic journals:


The five studies were published between 2002 and 2009. To pick just one, in 2007 Guéguen was the sole author of a study entitled “Bust Size and Hitchhiking: A field study“.

Good that they are being retracted now (16 to 23 years later), but embarrassing that they were published in the first place.
 
The party of Freedom and Small Government is now declaring what results science is allowed to give in the US. Sorta like Stalin decreed results during his term.
 
Great!
Now all we need is for each and every online post referencing those fake publications as valid to be removed. 😠
 
Article in Nature:


Earlier this year, computer scientist Guillaume Cabanac received a notification from Google Scholar that one of his publications had been cited in a paper published in the International Dental Journal. That was unexpected, because his research on spotting fabricated papers doesn’t typically intersect with dentistry. “I was very surprised to see that I couldn’t recognize my own reference,” says Cabanac, who is based at the University of Toulouse in France.

The title in the citation resembled that of a preprint he had posted in 2021 and never published formally, but the journal was listed as Nature and the DOI — the unique identifier assigned by publishers and preprint repositories — did not lead to the original preprint. “I got very concerned,” adds Cabanac, who immediately suspected that the citation had been hallucinated by artificial intelligence.

This is just one example of a rapidly growing problem. Surveys and related studies have shown that researchers are increasingly using large language models (LLMs) to help to conduct literature searches, write manuscripts and format bibliographies. And sometimes, these models generate non-existent academic references.
 
Should be something you could get AI to do...
AI invented the fake citations in the first place. You don't need AI to check a list of citations against an index of real papers.

I once asked ChatGPT to give me the citation to the first paper that discussed a particular method in quantum chemistry. It made up the citation. When I told ChatGPT that there was no such paper, it made up another one, and when I told it that paper didn't exist either, it made up another. not only never game the correct answer (which I knew in advance), it never gave a real paper.

This was less than two years ago. But I think AI has gotten better between then and now.
 
Last edited:
About the only relevant thing you may need AI for is to write the program. Then that program would not hallucinate.
Thought I'd see whether an AI could code something to do this. It could - a few problems along the way, mostly because of me insisting it all be contained in a single HTML file.

If you want to play with it copy the code into an editor that can save it as plain text, save the file as "citationchecker.html" (or any name you want) and drop the file into your browser, only works with mainly text PDFs.

HTML:
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<title>Darat -  Universal Citation‑Miner</title>

<style>
  :root {
    --bg-sage: #E8EDD9;
    --panel-bg: #FAF7F2;
    --text-main: #2F2F2F;
    --olive: #6B7A3A;
    --terracotta: #C46A4A;
    --line-soft: #D9DCC8;
    --badge-bg: #333333;
    --badge-text: #FFFFFF;
  }

  body {
    margin: 0;
    padding: 24px;
    background: var(--bg-sage);
    font-family: system-ui, -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif;
    color: var(--text-main);
  }

  h1, h2 {
    font-family: Georgia, "Times New Roman", serif;
    margin: 0 0 12px 0;
    color: var(--text-main);
  }

  h1 { font-size: 26px; letter-spacing: 0.03em; }
  h2 { font-size: 18px; letter-spacing: 0.02em; }

  .panel {
    background: var(--panel-bg);
    border-radius: 14px;
    padding: 18px 20px;
    margin-bottom: 20px;
    box-shadow: 0 10px 24px rgba(0,0,0,0.06);
    border: 1px solid rgba(0,0,0,0.03);
  }

  .pill {
    display: inline-flex;
    align-items: center;
    padding: 3px 10px;
    border-radius: 999px;
    font-size: 11px;
    background: #E1E6CF;
    color: var(--text-main);
  }

  button {
    border-radius: 999px;
    border: none;
    padding: 8px 18px;
    margin-left: 10px;
    font-size: 14px;
    font-weight: 500;
    cursor: pointer;
    background: var(--olive);
    color: #fff;
    transition: background 0.15s ease, transform 0.1s ease;
  }

  button:hover:not(:disabled) {
    background: #7C8B45;
    transform: translateY(-1px);
  }

  .accordion {
    border-radius: 10px;
    overflow: hidden;
    margin-bottom: 10px;
    border: 1px solid rgba(0,0,0,0.05);
  }

  .accordion-header {
    display: flex;
    justify-content: space-between;
    align-items: center;
    padding: 10px 14px;
    background: #F2F4E4;
    cursor: pointer;
  }

  .accordion-title {
    display: flex;
    align-items: center;
    gap: 8px;
    font-size: 14px;
    font-weight: 600;
    text-transform: uppercase;
  }

  .badge {
    display: inline-flex;
    align-items: center;
    justify-content: center;
    min-width: 24px;
    padding: 2px 8px;
    border-radius: 999px;
    font-size: 11px;
    font-weight: 600;
    background: var(--badge-bg);
    color: var(--badge-text);
  }

  .accordion-body {
    display: none;
    padding: 10px 14px 14px 14px;
    background: #FBFAF6;
  }

  .accordion-body.open { display: block; }

  table {
    width: 100%;
    border-collapse: collapse;
    margin-top: 6px;
    font-size: 13px;
  }

  th, td {
    padding: 10px 10px;
    border-bottom: 1px solid var(--line-soft);
    vertical-align: top;
  }

  th {
    text-align: left;
    font-size: 12px;
    text-transform: uppercase;
    color: #555;
  }

  tbody tr:nth-child(odd) { background: #FBFAF6; }
  tbody tr:nth-child(even) { background: #F5F6EB; }

  .status-ok { color: var(--olive); font-weight: 600; }
  .status-fail { color: var(--terracotta); font-weight: 600; }
  .confidence-high { color: var(--olive); font-weight: 600; }
  .confidence-medium { color: var(--terracotta); font-weight: 600; }

  a.doi-link {
    color: var(--olive);
    text-decoration: none;
    font-size: 12px;
  }
</style>
</head>
<body>

<div class="panel">
  <div style="display:flex; justify-content:space-between; align-items:flex-start; gap:12px; flex-wrap:wrap;">
    <div>
      <h1>Darat's Universal Citation‑Miner</h1>
      <div style="font-size:13px; max-width:520px;">
        Extract full references, inline citations, numeric citations, DOIs, footnotes, and more from any PDF.
      </div>
    </div>
    <div style="display:flex; flex-direction:column; align-items:flex-end; gap:6px;">
      <span class="pill">Worker‑free · Local‑friendly</span>
      <span class="pill">CrossRef + OpenAlex</span>
    </div>
  </div>

  <div style="margin-top:14px;">
    <label style="font-size:14px; font-weight:500;">
      Load PDF:
      <input type="file" id="pdfFile" accept="application/pdf">
    </label>
    <button id="extractBtn" disabled>Mine Citations</button>
  </div>

  <div id="status" style="margin-top:10px; font-size:13px;"></div>
  <pre id="error-log" style="margin-top:8px; font-size:12px; color:#C46A4A; white-space:pre-wrap; max-height:140px; overflow-y:auto;"></pre>
</div>

<div class="panel" id="results-panel" style="display:none;">
  <h2>Citations</h2>

  <div id="accordion-full" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>Full References</span>
        <span class="badge" id="badge-full">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-full"></div>
  </div>

  <div id="accordion-inline" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>Inline Citations</span>
        <span class="badge" id="badge-inline">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-inline"></div>
  </div>

  <div id="accordion-numeric" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>Numeric Citations</span>
        <span class="badge" id="badge-numeric">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-numeric"></div>
  </div>

  <div id="accordion-doi" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>DOIs</span>
        <span class="badge" id="badge-doi">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-doi"></div>
  </div>

  <div id="accordion-footnote" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>Footnotes</span>
        <span class="badge" id="badge-footnote">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-footnote"></div>
  </div>

  <div id="accordion-other" class="accordion">
    <div class="accordion-header">
      <div class="accordion-title">
        <span>Other / Unclassified</span>
        <span class="badge" id="badge-other">0</span>
      </div>
      <div>▼</div>
    </div>
    <div class="accordion-body" id="body-other"></div>
  </div>
</div>

<script src="https://cdnjs.cloudflare.com/ajax/libs/pdf.js/2.16.105/pdf.min.js"></script>
<script>
let pdfArrayBuffer = null;

const fileInput = document.getElementById("pdfFile");
const extractBtn = document.getElementById("extractBtn");
const statusEl = document.getElementById("status");
const errorLogEl = document.getElementById("error-log");
const resultsPanel = document.getElementById("results-panel");

const badgeFull = document.getElementById("badge-full");
const badgeInline = document.getElementById("badge-inline");
const badgeNumeric = document.getElementById("badge-numeric");
const badgeDoi = document.getElementById("badge-doi");
const badgeFootnote = document.getElementById("badge-footnote");
const badgeOther = document.getElementById("badge-other");

const bodyFull = document.getElementById("body-full");
const bodyInline = document.getElementById("body-inline");
const bodyNumeric = document.getElementById("body-numeric");
const bodyDoi = document.getElementById("body-doi");
const bodyFootnote = document.getElementById("body-footnote");
const bodyOther = document.getElementById("body-other");

function logError(msg, err) {
  errorLogEl.textContent += msg + (err ? " :: " + err : "") + "\n";
  console.error(msg, err);
}

/* -----------------------------------------
   ACCORDION BEHAVIOUR
------------------------------------------*/
document.querySelectorAll(".accordion-header").forEach(header => {
  header.addEventListener("click", () => {
    const body = header.nextElementSibling;
    body.classList.toggle("open");
  });
});

/* -----------------------------------------
   FILE LOADING
------------------------------------------*/
fileInput.addEventListener("change", async e => {
  const file = e.target.files[0];
  if (!file) return;
  try {
    pdfArrayBuffer = await file.arrayBuffer();
    extractBtn.disabled = false;
    statusEl.textContent = "PDF loaded.";
    errorLogEl.textContent = "";
    resultsPanel.style.display = "none";
  } catch (err) {
    logError("Error reading file", err);
  }
});

/* -----------------------------------------
   EXTRACTION WORKFLOW
------------------------------------------*/
extractBtn.addEventListener("click", async () => {
  if (!pdfArrayBuffer) return;

  extractBtn.disabled = true;
  statusEl.textContent = "Extracting text…";
  errorLogEl.textContent = "";
  resultsPanel.style.display = "none";

  try {
    const pdf = await pdfjsLib.getDocument({ data: pdfArrayBuffer }).promise;
    let pages = [];

    for (let i = 1; i <= pdf.numPages; i++) {
      const page = await pdf.getPage(i);
      const content = await page.getTextContent();
      const text = content.items.map(x => x.str).join(" ");
      pages.push(text);
    }

    const cleaned = cleanHeaders(pages);
    statusEl.textContent = "Mining citations…";

    const mined = mineCitations(cleaned);
    await verifyCitations(mined);

    renderGroups(mined);
    resultsPanel.style.display = "block";
    statusEl.textContent = "Done.";

  } catch (err) {
    logError("Extraction error", err);
    statusEl.textContent = "Error during extraction.";
  }

  extractBtn.disabled = false;
});

/* -----------------------------------------
   HEADER CLEANER
------------------------------------------*/
function cleanHeaders(pages) {
  const heads = pages.map(p => p.slice(0, 120).trim());
  const freq = {};

  heads.forEach(h => {
    if (h.length < 20) return;
    const key = h.slice(0, 60);
    freq[key] = (freq[key] || 0) + 1;
  });

  let header = null;
  let max = 0;
  for (const k in freq) {
    if (freq[k] > max && freq[k] >= 3) {
      max = freq[k];
      header = k;
    }
  }

  let text = pages.join("\n\n");

  if (header) {
    const safe = header.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
    const re = new RegExp(safe + ".*?(?=\\n|$)", "g");
    text = text.replace(re, "");
  }

  text = text.replace(/Page Proof.*?(?=\n|$)/gi, "");
  text = text.replace(/Compositor Name:.*?(?=\n|$)/gi, "");
  text = text.replace(/\bpage\s+\d+\b/gi, "");
  text = text.replace(/\b\d{1,2}\.\d{1,2}\.\d{2,4}\b/g, "");
  text = text.replace(/\b\d{1,2}:\d{2}(am|pm)\b/gi, "");

  return text.replace(/\s+/g, " ").trim();
}

/* -----------------------------------------
   CITATION MINER
------------------------------------------*/
function mineCitations(text) {
  const fullRefs = new Set();
  const inlineCits = new Set();
  const numericCits = new Set();
  const doiCits = new Set();
  const footnoteCits = new Set();
  const otherCits = new Set();

  function accept(str) {
    const s = str.trim();
    if (s.length < 25 || s.length > 350) return false;
    const bad = ["page proof", "compositor", "copyright", "printed in"];
    return !bad.some(w => s.toLowerCase().includes(w));
  }

  const doiRegex = /10\.\d{4,9}\/[-._;()/:A-Za-z0-9]+/g;
  let m;
  while ((m = doiRegex.exec(text)) !== null) doiCits.add(m[0]);

  const sentences = text.split(/(?<=[.!?])\s+(?=[A-Z])/g);

  const fullRefRegex = /[A-Z][a-z]+,\s+[A-Z](?:\.[A-Z]\.)?.*?\b(19|20)\d{2}.*?(?=[.!?])/g;
  sentences.forEach(s => {
    let mm;
    while ((mm = fullRefRegex.exec(s)) !== null) {
      if (accept(mm[0])) fullRefs.add(mm[0]);
    }
  });

  const inline1 = /\([A-Z][A-Za-z]+(?:\s+et al\.)?,\s*(19|20)\d{2}\)/g;
  const inline2 = /[A-Z][A-Za-z]+(?:\s+et al\.)?\s*\((19|20)\d{2}\)/g;
  sentences.forEach(s => {
    let mm;
    while ((mm = inline1.exec(s)) !== null) if (accept(mm[0])) inlineCits.add(mm[0]);
    while ((mm = inline2.exec(s)) !== null) if (accept(mm[0])) inlineCits.add(mm[0]);
  });

  const numRegex = /\[\d+(?:\s*,\s*\d+)*\]/g;
  sentences.forEach(s => {
    let mm;
    while ((mm = numRegex.exec(s)) !== null) if (accept(mm[0])) numericCits.add(mm[0]);
  });

  const footRegex = /\b\d{1,3}\s+[A-Z][a-z]+.*?\b(19|20)\d{2}.*?(?=[.!?])/g;
  sentences.forEach(s => {
    let mm;
    while ((mm = footRegex.exec(s)) !== null) if (accept(mm[0])) footnoteCits.add(mm[0]);
  });

  sentences.forEach(s => {
    const hasYear = /(19|20)\d{2}/.test(s);
    const hasName = /[A-Z][a-z]+,\s+[A-Z]\./.test(s);
    if (hasYear && hasName && accept(s)) {
      if (![...fullRefs].includes(s) && ![...footnoteCits].includes(s)) {
        otherCits.add(s);
      }
    }
  });

  return {
    full: [...fullRefs],
    inline: [...inlineCits],
    numeric: [...numericCits],
    doi: [...doiCits],
    footnote: [...footnoteCits],
    other: [...otherCits]
  };
}

/* -----------------------------------------
   VERIFICATION (CrossRef + OpenAlex)
------------------------------------------*/
async function verifyCitations(mined) {
  mined._verifiedFull = [];
  mined._verifiedDoi = [];

  for (const ref of mined.full) {
    const res = await checkReference(ref);
    mined._verifiedFull.push({ reference: ref, ...res });
  }

  for (const doi of mined.doi) {
    const res = await checkReference(doi);
    mined._verifiedDoi.push({ reference: doi, ...res });
  }
}

function extractPossibleDOI(ref) {
  const m = ref.match(/10\.\d{4,9}\/[-._;()/:A-Za-z0-9]+/);
  return m ? m[0] : null;
}

async function checkReference(ref) {
  let doi = extractPossibleDOI(ref);
  let sources = [];
  let score = 0;

  let cross = null;
  let alex = null;

  try {
    if (doi) {
      const r = await fetch("https://api.crossref.org/works/" + encodeURIComponent(doi));
      if (r.ok) cross = (await r.json()).message;
    }
    if (!cross) {
      const q = encodeURIComponent(ref.slice(0, 200));
      const r = await fetch("https://api.crossref.org/works?query.bibliographic=" + q + "&rows=1");
      if (r.ok) {
        const d = await r.json();
        if (d.message.items.length) cross = d.message.items[0];
      }
    }
  } catch {}

  if (cross) {
    sources.push("CrossRef");
    if (!doi && cross.DOI) doi = cross.DOI;
    score += 0.4;
  }

  try {
    const q = encodeURIComponent(ref.slice(0, 200));
    const r = await fetch("https://api.openalex.org/works?search=" + q + "&per-page=1");
    if (r.ok) {
      const d = await r.json();
      if (d.results.length) alex = d.results[0];
    }
  } catch {}

  if (alex) {
    sources.push("OpenAlex");
    if (!doi && alex.doi) doi = alex.doi.replace(/^https?:\/\/doi.org\//, "");
    score += 0.3;
  }

  const status = score >= 0.4 ? "ok" : "fail";
  const confidence = score >= 0.75 ? "High" : score >= 0.4 ? "Medium" : "Low";

  return { status, confidence, doi, sources, score };
}

/* -----------------------------------------
   RENDERING
------------------------------------------*/
function escapeHtml(str) {
  return str
    .replace(/&/g,"&amp;")
    .replace(/</g,"&lt;")
    .replace(/>/g,"&gt;");
}

function renderGroups(mined) {
  badgeFull.textContent = mined._verifiedFull.length;
  bodyFull.innerHTML = renderVerifiedTable(mined._verifiedFull);

  badgeInline.textContent = mined.inline.length;
  bodyInline.innerHTML = renderSimpleTable(mined.inline);

  badgeNumeric.textContent = mined.numeric.length;
  bodyNumeric.innerHTML = renderSimpleTable(mined.numeric);

  badgeDoi.textContent = mined._verifiedDoi.length;
  bodyDoi.innerHTML = renderVerifiedTable(mined._verifiedDoi);

  badgeFootnote.textContent = mined.footnote.length;
  bodyFootnote.innerHTML = renderSimpleTable(mined.footnote);

  badgeOther.textContent = mined.other.length;
  bodyOther.innerHTML = renderSimpleTable(mined.other);
}

function renderVerifiedTable(items) {
  if (!items.length) {
    return '<div style="font-size:12px; color:#666;">No items detected.</div>';
  }

  let html = `
    <table>
      <thead>
        <tr>
          <th>Status</th>
          <th>Confidence</th>
          <th>Reference</th>
          <th>DOI</th>
          <th>Sources</th>
        </tr>
      </thead>
      <tbody>
  `;

  items.forEach(r => {
    const statusClass =
      r.status === "ok" ? "status-ok" :
      r.status === "fail" ? "status-fail" : "";

    const confClass =
      r.confidence === "High" ? "confidence-high" :
      r.confidence === "Medium" ? "confidence-medium" : "";

    let doiHtml = "—";
    if (r.doi) {
      const url = "https://doi.org/" + r.doi;
      doiHtml = `<a class="doi-link" href="${url}" target="_blank">${escapeHtml(r.doi)}</a>`;
    }

    const src = r.sources && r.sources.length ? r.sources.join(" + ") : "—";

    html += `
      <tr>
        <td class="${statusClass}">${escapeHtml(r.status)}</td>
        <td class="${confClass}">${escapeHtml(r.confidence)}</td>
        <td>${escapeHtml(r.reference)}</td>
        <td>${doiHtml}</td>
        <td>${escapeHtml(src)}</td>
      </tr>
    `;
  });

  html += "</tbody></table>";
  return html;
}

function renderSimpleTable(list) {
  if (!list.length) {
    return '<div style="font-size:12px; color:#666;">No items detected.</div>';
  }

  let html = `
    <table>
      <thead><tr><th>Citation</th></tr></thead>
      <tbody>
  `;

  list.forEach(item => {
    html += `<tr><td>${escapeHtml(item)}</td></tr>`;
  });

  html += "</tbody></table>";
  return html;
}
</script>

</body>
</html>
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom