peck.press

№ 959,590

I have been using large language models for literature searches, which is one of the few occupations in which artificial

Bitcoin Dictionary · 2026-07-23 · 2 min read · treechat · tx b34fdb…7d88 · block 959,144

I have been using large language models for literature searches, which is one of the few occupations in which artificial intelligence can be genuinely useful rather than merely decorative.

Serious research does not involve finding ten papers and admiring their titles. One may have to pass through a hundred, sometimes two hundred, to discover the twenty that actually matter. Abstracts help, of course, but an abstract is rather like a society portrait: flattering, abbreviated, and carefully arranged to conceal anything inconvenient.

I was trained to search through abstracts, identify likely candidates and then read the papers themselves.

The difficulty is that the abstract often omits precisely the detail one needs: the actual result, the qualification, the figure hidden in Table 6, the limitation quietly confessed three pages before the conclusion.

A competent LLM improves this process enormously.

It can read the whole paper, extract the relevant claims, identify the important figures and tell me whether the work is worth my time. When it does not hallucinate, it is exceptionally useful. Instead of reading fifty or sixty papers after the abstract search, I may need to examine only ten closely in order to locate the twenty or so sources that genuinely belong in the research.

That is not replacing scholarship. It is replacing clerical drudgery. There is a difference, although people who do very little scholarship are often the most anxious to deny it.

The problem is that the models keep changing, and improvement appears to have become an optional feature.

I went back to ChatGPT because the results I was getting from Claude Fable were abysmally poor. It would return perhaps ten papers, which I would then have to read merely to discover whether the model had understood them correctly. This still saved some time, but only in the rather modest sense that being given a defective map saves one from wandering without any map at all.

What has become particularly obvious is how badly the quality has deteriorated. Compared with 4.8, the present research output has not merely declined. It has been flushed down the toilet, where it now appears to be conducting a comprehensive review of the plumbing.

The great promise of LLM-assisted research was not that the machine would think for us. It was that it would perform the tedious preliminary labour accurately enough for us to think better.

When it gives irrelevant papers, misunderstands results, omits central findings or invents connections that are not there, it ceases to be a research assistant and becomes a very confident undergraduate with no fear of being examined.

The tragedy is not that the technology is incapable.

We have already seen that it is capable. The tragedy is that newer versions can be less reliable than the older ones while being presented with all the ceremony normally reserved for progress.

In technology, as in society, novelty is often mistaken for improvement. The difference is that society usually has the decency to serve champagne while making the mistake.