Grok 3: Early Access Technical Review 1/ Recently, @karpathy with early access to Grok 3, @xai's latest large…

0 comments · 2025-02-18T06:14:35+00:00

@axeli (Web3Cryptos)

Grok 3: Early Access Technical Review 1/ Recently, @karpathy with early access to Grok 3, @xai's latest large language model, shared their first impressions and technical insights. In just a short session, they explored Grok 3’s reasoning capabilities, coding performance, search quality, and responses to classic LLM challenges. 2/ Reasoning and Test-Time Compute 🔥Grok 3 performed impressively on several reasoning and coding tasks, often matching or surpassing state-of-the-art models like OpenAI’s o1-pro and outperforming DeepSeek-R1 and Gemini 2.0 Flash Thinking. 3/ 🔥Coding Challenge on Settlers of Catan Board ✅ The @karpathy asked Grok 3 to generate HTML for a Settlers of Catan-style hex grid with an adjustable slider for the number of rings. Grok 3 solved this prompt—a challenge that only OpenAI’s o1-pro has reliably passed. Competing models like DeepSeek-R1, Gemini 2.0 Flash Thinking, and Claude failed. 4/ Unicode Puzzle: Emoji Mystery ❌ @karpathy tested Grok 3’s ability to decode a hidden message embedded in Unicode variation selectors, providing hints in Rust code. Grok 3 failed, though DeepSeek-R1 previously managed a partial solution. 5/ Tic-Tac-Toe Reasoning ✅ 🔥Grok 3 solved multiple tricky tic-tac-toe boards using clear reasoning, outperforming many state-of-the-art models. However, when asked to generate tricky boards, it failed—producing nonsensical layouts. Notably, OpenAI’s o1-pro failed here as well. 6/ Math and Estimation- GPT-2 Training FLOPs ✅ 🔥In a demanding reasoning test, the Andrej asked Grok 3 to estimate the number of FLOPs required to train GPT-2 based on the GPT-2 paper. Grok 3 succeeded using “Thinking” mode, correctly combining knowledge, lookup skills, and estimation. OpenAI’s o1-pro failed this challenge. 7/ Riemann Hypothesis Reasoning ✅ 🔥When challenged with the Riemann Hypothesis, Grok 3 attempted a step-by-step reasoning process rather than simply stating it was an unsolved problem. The Andrej ended the response but noted that Grok 3 demonstrated persistence and genuine reasoning effort, unlike many models that immediately dismiss it as unsolvable. 8/ DeepSearch ( Strong Research Capabilities, Needs Polishing) Grok 3 features a search-augmented mode called DeepSearch, similar to Perplexity’s “DeepResearch” or OpenAI’s “Deep Research.” The Karpathy found it capable on current events but noticed some limitations in search coverage and source reliability. 9/ ✅ Strong Performance on Trending and Factual Queries 🔥“What’s up with the upcoming Apple launch?” → Accurate, well-referenced response. “Why is Palantir stock surging recently?” → Clear, well-sourced explanation. “Where was White Lotus Season 3 filmed?” → Correct location and production details. “What toothpaste does Bryan Johnson use?” → Correct answer from available sources. 10/ ❌ Weak Performance on Niche or Time-Sensitive Queries “Singles Inferno Season 4 cast—where are they now?” → Provided incorrect information. “What speech-to-text tool is Simon Willison using?” → Couldn’t retrieve relevant results. 11/ ⚠️ DeepSearch Issues Observed 🦴Citation Errors- Sometimes referenced URLs that did not exist. 🦴X/Twitter Coverage Gap-Underutilized @x unless explicitly prompted. 🦴Knowledge Gaps- Failed to list xAI as a major LLM lab in a funding and employee count comparison, which suggests blind spots in its knowledge graph. Overall, the @karpathy found DeepSearch comparable to Perplexity’s DeepResearch but not yet on par with OpenAI’s Deep Research, which offers more thorough citations and coverage. 12/ LLM Gotchas and Reasoning Tests 🔥Karpathy tested Grok 3 with several classic LLM challenges-queries designed to be easy for humans but often tricky for language models. ✅ Counting Letters: “How many ‘r’ are in ‘strawberry’?” → Correct (3). ✅ Number Comparison: “Which is larger: 9.11 or 9.9?” → Initially failed but succeeded with “Thinking” mode. ✅ Logic Puzzle: “Sally has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?” → Correct (2). Many models fail this one. ❌ Joke Quality Check: Generated a basic pun: “Why did the chicken join a band? Because it had the drumsticks and wanted to be a cluck-star!” The humor was typical of LLMs—surface-level and uninspired. Areas for Improvement 💯Despite strong performance in reasoning and search tasks, the Karpathy identified several areas where Grok 3 fell short. ✍️Visual Reasoning (SVG Generation)- Failed the “Pelican on a Bicycle” SVG generation test, a known challenge for spatial reasoning. Claude outperformed here. ✍️Creative Humor-Joke generation was basic and lacked creativity, a common weak spot for LLMs. ✍️Handling Ethical Scenarios-The model was overly cautious, producing a long, noncommittal essay when asked a nuanced ethics question🤔 ✍️Search Coverage & Citations- Needs more reliable citations and broader coverage of niche queries, especially from X/Twitter 🤔 ☺️Overall Technical Impression 🔥In just two hours of testing by @karpathy, Grok 3 demonstrated state-of-the-art reasoning capabilities and strong research performance, especially with its “Thinking” mode. 🔥Its performance on coding tasks, math reasoning, and logic puzzles rivals OpenAI’s o1-pro and outpaces competitors like DeepSeek-R1 and Gemini 2.0 Flash Thinking. 🔥The DeepSearch feature is promising but needs refinement to match OpenAI’s Deep Research in coverage and citation reliability. Key Strengths ✅ Outstanding on complex coding and reasoning tasks. ✅ Excellent performance on math-heavy logic problems. ✅ Strong search augmentation for recent events and trends. Key Weaknesses ❌ Limited creativity in humor and storytelling. ❌ Poor performance on spatial reasoning tasks (e.g., SVG creation). ❌ Inconsistent citation quality and search coverage. Final Thoughts 🎯For a model developed in under a year, Grok 3 is a remarkable achievement from xAI, standing shoulder-to-shoulder with some of the best models from OpenAI while clearly surpassing others like DeepSeek-R1 and Gemini 2.0 in reasoning and logic tasks. 🎯The DeepSearch feature shows strong potential but could use more polish. As the model continues to evolve, Grok 3’s rapid progress signals that xAI is quickly becoming a serious contender in the LLM space. 🎯This first look suggests that Grok 3 will be a valuable addition to the LLM ecosystem, particularly for coding, math, and research tasks. It's a model to watch closely as it continues to develop. \~ follow @Web3Aible
Profile · Hey · Permalink