Jaak Henno, Tallinn University of Technology, Estonia
Hannu Jaakkola, Tampere University of Technology, Pori
Jukka Mäkelä, University of Lapland, Rovaniemi, Finland
30-40% of results produced by LLMs are erraneous
It is impossible to know, which ones
Error rates in new LLM-s are not decreasing, but increasing
Improper/vague questions
Unexplained ("black-box") answers
| The common 4 data points can produce very different curves
Values for future (forward approximation) differ more than twice |
y1(5) = 18; y2(5) = 39.6 |
neural net hallucinations |
Interpolation from chatGPT (order was not given) |
The road created by chatGPT |
The real ice road |
The road created by chatGPT |
The road created by Sonnet 4.6 |
Create web page showing map rectangle (60.4, 24.0); (59.4,30.5); indicate on map Tallinn, Helsinki, St.Petersburg, Primorsk, Ust-Luga and cable zone lat=24.6, lat=25.2; record time, coordinates and name of all tankers and cargo ships when they enter the cable zone.
Show real ships, do not animate
The web-page looked like a real good app - ships moved and created warning when they entered the cable zone
Only after following it for a long time (or inspecting the www-page Javascript) it appeared, that all ships were animations, "a Potjomkin village", but Gemini did not mention this.
| Generate a Javascript program to draw on a html-page a centered n x n grid of 50 x 50 pixel visible cells with light border; in every row and column are randomly integers 1,2,...,n-m (m < n) and m empty cells filled with gray | ![]() |
All these 'success stories' have one thing in common - they are not reproducible
This slope is growing. Crawlers from LLMs are constantly adding whatever they found to LLMs training corpuses, so new LLMs are trained on worse data
Error rate of new LLMs e.g. on coding tasks and multi-step reasoning is often 39%–60%
Give to LLMs controllable tasks:
not "calculate ...", but "give me a program to calculate...", but LLM-s can cheat even now - program results are not based on anything (ice road, cable monitor))
Formulate your tasks very carefully and as fully as possible; write the task and re-read it several times before presenting to LLM
Students should always remember the most often repeated guideline of LLM-s:
Check! Check! and then CHECK ONCE MORE!
EEG scans of students using ChatGPT show reduced brain activity
cognitive disempowerment - users can not recall or explain content that was generated by an LLM
skill atrophy and 'homogenized' thinking, less original, predictable ideas
LLMs are flooding the web with unoriginal, low-quality text, code, images. Even in science is growing production of 'text slope':
|
fake citations uncovered by audit of 2.5 million biomedical science papers |
|
Moderately helpful for specific tasks, but still need heavy supervision - 31%
Game-changer — they're already replacing big chunks of my routine - 17%
Actively annoying — more hallucinations and security risks than value - 15%
I refuse to use them on principle - 11%
Haven't tried them seriously yet - 10%