Steven, a graduate in technology studies and longtime New World Notes reader, has been working with a non-profit org to help develop a system for tracking infectious disease, and his higher-ups have been encouraging the team to use ChatGPT to build out the data base.

The trouble is, when Steven asked ChatGPT for knowledge of recent outbreaks of monkeypox, he got this result:

ChatGPT hallucination monkeypox

ChatGPT fails to mention that there is also a current outbreak in Africa. Nor mentions the outbreak last year. (Nor that it's now called "mpox".)

Again, this answer was going to be put into a database of infectious diseases for use by NGOs and government bodies, before a human (i.e. Steven) caught the error.

"[I'm] not quite sure this is the slam dunk my boss thinks this is," as he puts it about ChatGPT.

To be sure, the program does include some warnings about the possibility of hallucinations (i.e. incorrect output) and that version 3 — the one this non-profit project is using — is only trained on data up to 2021. But Steven is concerned that the average user may not notice these qualifiers.

"[ChatGPT has a] quick dialogue box when you first open the page, but who reads that?"

And again, Steven's superior is the one who ordered staff to develop the database with ChatGPT!

"[The boss] cited a couple Insider articles," then got excited by the hype cycle. And this isn't using ChatGPT's output for entertainment value, but in a serious application: 

"We are trying to develop a number of crowdsourced surveillance systems for infectious disease in order to essentially create 'an apples to apples information pipeline' for government, enterprise, and the civic population all working together in concert."

This isn't to say ChatGPT is completely useless in this context. 

"Overall I was impressed with website scraped data it was able to pull," as Steven puts it. "But my problem is that it wasn't giving explicit sources, which you kinda need, especially around science. [That] and the 'bias' for as recent as possible data."

So ChatGPT is useful — but with qualifiers:

"If you need information quick that can be scraped, and don’t need a source, it’s good. But anything that needs a source or up to date info is going to be problem. To me it’s not so much intelligence as it is a fast scraper of large sets of data. Understand that in the advent of mega data that can be useful, but for small scale applications, it’s not really all that much help."

ChatGPT will need serious improvements before being more than that.

"Essentially what the next stage AI researchers need to do is make a 'living' data set because you can’t artificially choose a set date and then call this an intelligence. At best ChatGPT is text-based analytic software."

This has been my own personal impression with ChatGPT. I recently experimented with it on a writing project (not this blog or my book, of course!), and the output was so mediocre or wrong, it would take more time correcting and polishing its output, than just writing the whole thing myself. And forget about asking it for help with a topic you actually know well.

ChatGPT is fun and valuable in limited use cases. But with over 100 million people now using it, my concern is there are many millions who approach it with highly inflated expectations, literally risking people's lives by applying ChatGPT's output to situations for which it is simply not designed for.

To judge by my Twitter feed, where every other tweet treats ChatGPT as if it is an omniscient deity, that can't be said loudly enough.  

"The medical techno-futurists are seeing AI as the future of medicine," as Steven puts it, "but the thing is it needs to be guided by human clinicians. But we all know that one doctor who could give a shit about guidance and quality care, and just accept the AI diagnosis when it is the exact wrong thing."

Posted in

8 responses to “ChatGPT Is Not Reliable in Serious Use Cases – Here’s Yet Another Example”

  1. grace Avatar
    grace

    ChatGPT does not mention the current outbreak in Africa nor the outbreak last year because its training data (I am not sure for which version) only goes up to 2021.

    Like

  2. grace Avatar
    grace

    .. so in other words, it is not an AI issue but a user (human) issue. Where the facts are important because it will be relied on, the human needs to check / verify the ‘facts’ reproduced by the AI from somewhere else are indeed facts. Just as you won’t trust all the information on the web to put into a medical database without some level of verification or fact-check, why would you do differently (ie. trust the AI) because a software was a medium between the human and the web?

    Like

  3. Steven Avatar
    Steven

    Grace -I don’t know how many people are going to catch that because before I used it I thought it was up to date, but like I mentioned, then how can we functionally use ChatGPT for time-sensitive industrial use? I didn’t even know that ChatGPT is just “research” and not a ready-to-go piece of software that is delivering all the promises that its hype people are claiming it to be, which is why my boss wants us to use it. Remember AI isn’t just marketed as a tech industry-only application, but as widespread web3 technology that will change how we live. If I have to essentially keep fact-checking the machine, why bother with the machine at all? It does nothing that humans already do, and I don’t know where that data is ultimately coming from.

    Like

  4. Luther Weymann Avatar
    Luther Weymann

    I remain unconcerned about the rise of the machines. I know what to do in a data center server room and how to shut down the database access points and the internet connections. I wonder what ChatGPT would say about that disconnect?

    Like

  5. seph Avatar
    seph

    I mean, it says right on the ChatGPT interface: “ChatGPT may produce inaccurate information about people, places, or facts”
    We can and should follow up with feedback and corrections to get the correct info we want. To me, that’s intuitive and fine. We’ve accepted having to Google more than once before giving up and ask Siri questions different ways before giving up, I don’t see the problem with having to guide ChatGPT a little bit ‘specially since we’re all used to doing it with search queries and voice assistants.
    In the case of recent history that needs to be accurate and we can’t explain it, ChatGPT is releasing plugins https://openai.com/blog/chatgpt-plugins that’ll allow it to search the web live, travel info live from sources like Expedia, etc. That’ll make it up to the second aware depending on what plugin you’re using.

    Like

  6. Nadeja Avatar
    Nadeja

    Eh… It’s a language model, not an oracle. Especially GPT-3.5 that was used there.
    Some people have wrong expectations, indeed, they think that large language models (LLMs) store information like databases and they process that like computers; instead we are taking of artificial neural networks here and they don’t work that way. They are inspired by biological neural network. A computer can run these models, like it can run a SNES emulator (it’s not an emulation, though, at most a rough simulation); but the “simulations”, the models, are not the same as the computer. They can and do mistakes. They “hallucinate” and shouldn’t be trusted.
    If you want to use a LLM to retrieve current information, then use one that can search.
    “ChatGPT will need serious improvements”. Already done: GPT-4. And GPT-4 with plugins. And GPT-4
    You can get GPT-4 that can search for information:
    – as ChatGPT when you pay the subscription for it and the plugins
    – as Bing Chat, that uses the search engine
    This is what Bing Chat does when you ask: “was there an outbreak of monkeypox?”
    https://i.ibb.co/whCZyqW/Bing-Chat-MPox.webp
    – It used the search engine to find updated information on the web,
    – then wrote the answer
    – it linked the sources.
    3 screenshots, because Bing Chat has 3 modalities. The plum colored one is the “creative” mode: more verbose, but it tends to hallucinate more (in this case, the confirmed cases and deaths correspond with what is written in Wikipedia). However, it’s useful in many ways. E.g.: I used it when I needed an idea for a name.

    Like

  7. John Smith Avatar
    John Smith

    I don’t get the feeling Luther (above) has seen Lawnmower Man

    Like

  8. Luther Weymann Avatar
    Luther Weymann

    John, not only have I seen Lawnmower Man but I used to have 11 valid patents on computer networks in neural environments filed between January of 2000 and April of 2002. Only two of those patents are relevant anymore. I owned that software company that paid for those eleven patents and nowadays as an old elderly retired man, I just wish I had the over half-million dollars in legal fees and research back in my account because tech has moved on and only two have any value anymore. I used to give NorCal conference talks on the future of computer neural networks back in the early 2000s. Things have changed, but not that much. By the way, Peter Tippett invented the basis of the computer neural networking when he wrote the original code that became Norton antivirus. Most people don’t like history, they think stuff just happened yesterday.

    Like

Leave a comment