Here is how AI measures time.
Here is how AI measures time.
If you keep sharing things like this, Iām going to end up with a flat spot on my forehead. ![]()
Judging by your avatar, it might take a while ![]()
Five months ago i got a Pixel Watch 4. I thought it was cool to tap my watch and ask Gemini a question.
After seeing these videos i might ask Gemini now about the weather, but for important things like a medical condition i search and do the vetting myself.
Could be. That character cracks me up, though. ![]()
Iāve learned that certain digital assistants do better with certain things than others. Amazon Alexaās weather information often differs quite a bit from Google Assistantās, which seems to agree more with what NWS says.
This doesnāt happen with some [good] AIs. Good AIs are starting use tool calls in appropriate situations or use other methods such as timestamps.
Math, counting, and spelling weaknesses are easily overcome using Python, Wolfram Alpha tools and Chain of Thought reasoning models. Time often uses timestamps and other tool calls now.
I rarely use OpenAI directly (I do use it indirectly sometimes through partnerships), so I am not quite as familiar with them, but lots of others have addressed all these common token predictor averaging shortfalls on their higher reasoning models.
There is a huge difference between basic consumer flash models vs the advanced reasoning agentic models and harnesses nowadays and it is hard to lump them all as the same thing. My self-created multi-agent swarm doesnāt have any of these issues and I have multiple models do advanced reasoning peer reviews of everything important and include and verify sources and accuracy. It actually works pretty well to have a sort of advisory board of peer reviews with advanced reasoning and logic. I still check things myself, but it is easy to review sources that have already been vetted multiple times from different perspectives so I am not wasting as much time when I review and approve/commit things.
But it is funny to see people using the small no reasoning flash models without tool call capabilities show the weaknesses of relying solely on token predictor averages in certain use cases. Because letās face it, the AVERAGE person only has an IQ of 100 by definition, and if we are relying on the average thinking and reasoning capability to do something extremely advanced and important, we could be in trouble. I used to work customer service through College for multiple fortune 500 companies, and in the middle of that experience I used to tell people that it taught me that most āpeople [on average] were basically morons with occasional bursts of intelligenceā (try explaining bills or math to the average person who canāt do basic addition but insists their bill is wrong)ā¦and while I have adjusted my view and understanding in the subsequent years (least of which being a sampling bias), I can understand the concern about relying on averages. Yes, AI is trained on advanced PhD and scientific info, but itās also trained on the dumbest things ever said with the most ignorant fringe forums in complete opposition to that, and it may conclude with a prediction on the average of that if not prompted in the right way.
I still enjoy the non-reasoning flash model memes like this
even if it doesnāt really apply to my swarm which uses a ton of agentic harnesses and tool calls for verification purposes. In some ways I hope the frontier model companies keep allowing these kind of issues.
Allows me to keep an edge a little longer.
I generally plant my garden or fertilize my lawn around the NWS forecast. For the Oregon area, I find NWS to be the most accurate.
Here is the link if anyone wants to check out the National Weather Service (NWS) for your area of the US.
I understand your point. Reminds me back in the 80s when I had a Corvette. I had a little Yugo pull up next to me a was gunning its engine.
There are big differences in car models and AI models.
āThe RAM shortage could last yearsā