They will pretty soon be able to fact check everything they say when they gain real time internet connectivity.
But for now yeah you're right. A year-two from now this won't be true anymore.
Being connected to the internet may make their information more up to date but they won’t stop being confidently incorrect. These language models are amazing at token prediction but it’s a mistake to believe that they’re actually intelligent.
Give it an inane prompt to better see the difference, like "write me an essay about camels and their value in cybersecurity" which gives us the answer
"The camel is a valuable animal in the cybersecurity field due to its ability to store large amounts of data in its hump. Camels are able to cross vast distances with little water, which makes them ideal for carrying large amounts of data across networks. They are also very sturdy animals, able to withstand harsh conditions and even attacks from cyber criminals. In addition, their long eyelashes protect their eyes from sand and dust, making them perfect for working in dusty environments such as data centers."
Well it did what you asked him to do , you ordered it to write an essay about camels and their value in cybersecurity. It doesn't understand if you're joking or not or what the purpose of the whole thing.
I asked it this:
Are camels important to cyber security?
Answer:
"No, camels are not typically considered to be important for cyber security. Camels are domesticated mammals that are well-adapted to desert environments and are often used as a means of transportation or for their milk, meat, and hides. Cyber security, on the other hand, is the practice of protecting computer systems, networks, and data from unauthorized access, use, disclosure, disruption, modification, or destruction. The two are not related."
A human would say that’s ridiculous and tell you why. Google would give you a link to Apache Camel. ChatGPT tells you about how great a camels humps are for data storage.
That’s preposterous. An intelligent agent, human, mineral, or otherwise, would respond that this is a ridiculous idea and ideally explain the reasons that such is the case. Imagine you are a student and you asked this amazing AI sone question of similar if mildly ridiculed and turn imagine the student didn’t already know the answer. Would you think this kind of response would be an example of an intelligent AI?
If it cannot deal with such things without being prompted in such a way that the prompter knows the answer already, how could it deal with complex legal situations with actually intelligent adversaries?
This is overly optimistic. For one, fact checking is much harder than you think it is. Aside from that, there are also many additional problems with AI legal representation, such as lack of body language cues, inability to formulate a coherent legal strategy, and bad logical leaps. We're nowhere near to solving those problems.
AI hallucinations are going to be the new database query injection. Saying that real time internet connected fact checking will solve that is every bit as naive as thinking the invention of higher level database abstractions like an ORM will solve trivially injectable code.
We can't even make live fact checking work with humans at the wheel. Legacy code bases are so prolific and terrible we're staring down the barrel of a second major industry crises for parsing dates past 2037, but sure LLM's are totally going to get implemented securely and updated regularly unlike all the other software in the world.
I'd also argue that "hallucination" is, at least in some form, pretty commonplace in courtrooms. Neither lawyers' nor judges' memories are foolproof and eyewitness studies show that humans don't even realise how much stuff their brain makes up on the spot to fill blanks. If nothing else, I expect AI to raise awareness for human flaws in the current system.
That the legal system has flaws isn't a good argument for allowing those flaws to become automated. If we're going to automate a task, we should expect it to better, not worse or just as bad (at this stage it would definitely be worse).