Herbie Creative

Sherbert. Serving up the daily AI scoop.

The scoop, Sep 11, 2026.

xAI targets Grok 4.7 for launch on September 12. The 2.1 trillion parameter model incorporates SpaceX data and lands alongside new tools that include a hosted MCP server for pulling docs into Grok Build, Cursor, and Zed, native image-to-video on grok-imagine-video-1.5, and Grok Build 1.0.25 with vim-mode navigation plus silent successful hooks. OpenAI launched GPT-6 Astra, its new flagship that claims top scores in software engineering, science, and cybersecurity with an initial rollout to select businesses.

  • 153 stories
  • 34 market
  • 12 open-weight models

Market

The facts up top. Quotes, chatter, and the full PTR stay folded.

  • -2.28% basket 1d
  • Leader AAPL +3.56%
  • NVDA vs QQQ 21d +1.87%

Pelosi PTR filed 2026-08-21. BE +28.31% since file. INTC +11.38% since file.

Copy-trade Hot PelosiTracker & Autopilot actively posting new pol trades w/ strong engagement

Quotes, chatter, and full PTR
NameLast1d5d21d
NVDA Nvidia218.36-2.26%-2.59%+0.51%
AMD AMD503.60-3.36%+10.18%+6.17%
AVGO Broadcom360.83-0.97%-1.75%-13.28%
TSM TSMC428.03-1.68%+3.02%+1.41%
ASML ASML1,687.43-2.43%+0.30%-6.22%
SMCI Super Micro37.38-3.98%+1.03%+18.29%
INTC Intel100.32-5.57%+11.40%+2.67%
MU Micron977.41-4.90%+2.23%+12.54%
ARM Arm254.18-3.80%+8.23%-5.48%
CRWV CoreWeave89.12-6.13%+10.12%-1.33%
MSFT Microsoft492.44+0.16%-0.88%-2.07%
GOOGL Alphabet332.60+0.59%-1.28%-3.20%
AMZN Amazon251.89-0.20%-1.21%-7.49%
META Meta644.38-1.42%+8.69%+7.55%
AAPL Apple326.57+3.56%+0.50%+7.10%
TSLA Tesla363.56-1.16%+1.83%+9.24%
ORCL Oracle152.94-5.38%+4.93%+5.13%
PLTR Palantir165.86-2.16%-2.12%-5.19%

Source: delayed Yahoo Finance, Stooq if Yahoo blanks. Do not treat this as a ticket.

  1. xAI Targets Grok 4.7 Launch September 12 xAI
  2. xAI Docs MCP Server now available for Grok Build, Cursor, and Zed xAI
  3. Image-to-Video support documented for grok-imagine-video-1.5 with native 1080p xAI
  4. Grok Build 1.0.25 released with silent successful hooks, vim-mode navigation, and workflow controls xAI
  5. grok-imagine-image-quality model retirement scheduled for November 2, 2026 xAI
  6. CoreWeave (CRWV) -6.13% today, +10.12% over 5 sessions CRWV
  7. AI and tech tape: delayed closes and 21-day math NVDA
  8. AMD (AMD) -3.36% today, +10.18% over 5 sessions AMD

Pelosi PTR

The house files late. People still copy it. That bid can lift the names they still hold. Maybe a bubble. Short game: ride the tide, sell when you are up.

Pelosi household filed a Periodic Transaction Report on 2026-08-21 filed 2026-08-21

AI and tech purchases on this filing. The ones the copy-trade crowd is chasing.

NameSideWhatAmountTradedLagLastSince tradeSince file
BE Bloom EnergyPurchaseStock$1,000,001 - $5,000,0002026-07-2428d258.49+39.81%+28.31%
Data-center power. The AI electricity trade. Purchased 10,000 shares.
BE Bloom EnergyPurchaseOptions$1,000,001 - $5,000,0002026-07-2428d258.49+39.81%+28.31%
Data-center power. The AI electricity trade. Purchased 100 call options with a strike price of $100 and an expiration date of 6/17/27.
BE Bloom EnergyPurchaseStock$500,001 - $1,000,0002026-07-2824d258.49+54.93%+28.31%
Data-center power. The AI electricity trade. Purchased 5,000 shares.
BE Bloom EnergyPurchaseOptions$500,001 - $1,000,0002026-07-2824d258.49+54.93%+28.31%
Data-center power. The AI electricity trade. Purchased 100 call options with a strike price of $100 and an expiration date of 6/17/27.
INTC IntelPurchaseOptions$250,001 - $500,0002026-07-2428d100.32+8.67%+11.38%
Chip. Foundry and AI hardware story. Purchased 50 call options with a strike price of $50 and an expiration date of 6/17/27.
INTC IntelPurchaseStock$500,001 - $1,000,0002026-07-2428d100.32+8.67%+11.38%
Chip. Foundry and AI hardware story. Purchased 10,000 shares.

Also on the filing, not a tape name.

  • REOF XXV, LLC · Purchase · $500,001 - $1,000,000 · Additional investment in LLC which is acquiring and restoring a luxury hotel property in San Francisco, CA.

Parsed from the official House PDF when a new doc_id hits the Clerk index. Amounts are ranges. Many rows are spouse (SP). Delayed quotes. Not a ticket.

The Scoop

What I clipped from today's AI news.

Grok pulse Market xAI

xAI Targets Grok 4.7 Launch September 12

xAI preparing Grok 4.7 release with 2.1T parameters incorporating SpaceX data; follows pattern of frequent model releases including recent Grok Bot enterprise availability and Grok 4.6 deployments to Vertex AI and Bedrock.

just now
Grok pulse: market Market NVDA

DOJ probes Nvidia licensing deal with AI startup Groq

U.S. Justice Department investigating whether Nvidia structured its $17 billion non-exclusive licensing deal with Groq (announced last year) to avoid antitrust scrutiny. Probe opened shortly after December announcement; Nvidia received formal request for information. Reported Sept 10, 2026.

just now
Grok pulse: market Market Anthropic

Unusual Whales on Anthropic researcher AI extinction warning

Posted Sept 10, 2026: Anthropic researcher Jacob Coxon, who resigned Tuesday, warned on X that people building AI believe it could kill everyone by the end of the decade. Links to full story on unusualwhales.com.

just now
Grok pulse: market Market Pelosi Autopilot Pelosi PTR

Autopilot Pelosi Tracker platform remains active

Platform for automatically copying politician trades including Pelosi reports over $1.1B invested by 180k users. Associated X accounts (@pelositracker) posted tracker alerts multiple times on Sept 10, 2026. Trend has not faded.

just now
Grok pulse

OpenAI Launches GPT-6 Astra Frontier Model

OpenAI releases GPT-6 Astra claiming it is the world's most intelligent model with leads in software engineering, science, and cybersecurity; initial limited release to select businesses and cybersecurity partners; priced comparably to leading Anthropic models but more efficient per task.

just now
Grok pulse

Anthropic Releases September 2026 Threat Intelligence Report

Detected large-scale distillation campaigns targeting Claude Opus models; Alibaba-affiliated operators ran campaign with over 151 million exchanges extracting chain-of-thought reasoning; additional campaigns attributed to Moonshot, DeepSeek, and Zhipu; documented Claude misuse for cyberattacks on over 20 organizations and building mass-surveillance platform in Mali.

just now
Grok pulse

Mistral Raises $3.5 Billion in Series D

Mistral completes record 3 billion euro (~$3.5B) funding round valuing company at 21 billion euros (~$24B); proceeds fund frontier research, compute infrastructure expansion, and commercial growth to compete on scale with US and Chinese labs.

just now
Grok pulse

Meta Releases Muse Spark 1.3 Model

Meta launches most powerful AI model to date; available to developers via paid API access immediately and rolling out soon to Instagram, Facebook, and Meta AI users; chief AI officer states capabilities now edging closer to OpenAI and Anthropic frontier models.

just now
Grok pulse

Anthropic Expands Project Glasswing

Anthropic expands Glasswing initiative ahead of cybersecurity vulnerabilities from Mythos Preview frontier model; partners use model to scan codebases for flaws; coordinates with US government and security industry.

just now
Grok pulse

Black Duck Joins Anthropic Project Glasswing

Black Duck integrates Mythos model into its application security portfolio for AI-accelerated vulnerability discovery combined with deterministic testing, remediation, and governance workflows; announced September 8.

just now
WIRED AI

Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online

Clearview AI tested InquiryIQ, a prototype tool powered by xAI's Grok model, to help law enforcement surface social accounts and associates tied to identified individuals. The tool raises surveillance and privacy concerns as it automates the aggregation of public data into actionable intelligence for police.

19h ago
Hacker News Paid tool

OpenAI Agents API

OpenAI launched an Agents API for building and running autonomous agents with built-in tool use and memory. Pricing follows standard API rates.

9h ago
Hacker News Free tier

DeepSeek v4.1 Flash

DeepSeek released v4.1 Flash, a smaller and faster variant of its reasoning model. Available via API with competitive pricing.

23h ago
OpenAI News Paid tool

Introducing ChatGPT for Financial Services

OpenAI shipped ChatGPT for Financial Services, bundling GPT-6 Astra with built-in financial data and tools for research and modeling. It is a vertical product aimed at the financial sector.

22h ago
WIRED AI

Why So Many AI Researchers Think the Machines Could Kill Everyone

Researchers at major AI labs are increasingly concerned about existential risk from AI systems, citing rapid capability gains, recursive self-improvement, and autonomous agent swarms as genuine sources of worry. The concern is real enough to shape how labs approach safety and deployment.

just now
TechCrunch AI

AI agents are flooding public services with new requests

AI agents are filing benefit claims and public service requests at scale, but most are legitimate. The flood is raising questions about how public systems handle automated submissions and whether oversight can keep pace.

14h ago
MIT Technology Review AI

Powering AI is an architecture problem

Power grid failures in data center clusters are becoming a bottleneck for AI infrastructure. A single fault in Virginia knocked 3 gigawatts offline, exposing how concentrated compute density creates single points of failure that can cripple AI services.

18h ago
The Decoder

OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT

OpenAI is releasing the Agents API as a public beta. It lets developers build cloud agents that run autonomously for hours, execute code, and hand off tasks to sub-agents. There are no extra fees beyond token usage. Cloudflare, Vercel, and Oracle offer additional sandbox environments. The article OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT appeared first on The Decoder .

just now
The Decoder

Former Deepmind PR staffer says the lab once banned public discussion of AI extinction risk

A former Google DeepMind spokesperson says talk of AI-driven human extinction was "external communication about the possibility of human extinction was not permitted, by anyone, at any level of the organization." Internally, the team knew AI alignment was not solved, according to Vishal Maini. The article Former Deepmind PR staffer says the lab once banned public discussion of AI extinction risk appeared first on The Decoder .

12h ago
OpenAI News

Now everyone can put data to work

Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.

14h ago
The Decoder

Claude Fable 5.1's language is less "load-bearing" than its predecessor's

Arena.ai analyzed how Claude's writing changed from Fable 5 to Fable 5.1 across tens of thousands of benchmark responses. Fable 5.1 writes more matter-of-fact but also more verbose. The article Claude Fable 5.1's language is less "load-bearing" than its predecessor's appeared first on The Decoder .

13h ago
Hacker News

Neijuan

24 points, 8 comments

just now
WIRED AI

Is AI Actually Going to Kill Us All?

This week on “Uncanny Valley,” we dig into a former Anthropic researcher’s AI doomsday warning, the latest upgrades from Apple’s event, and the census report that claimed Trump won the 2020 election.

8h ago
The Verge AI

Meta’s Muse AI works and creeps me out

Meta has launched its new Muse assistant, marking the company's first real foray into AI-powered productivity tools. The company says its AI agent can "take the busywork off your plate" by helping you with online shopping, emails, trip-planning, and more. I decided to try out the new tool and see how well it performed - […]

14h ago
The Verge AI

Mathematicians want proof OpenAI didn’t use their work

Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company's models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and "dishonest" behavior and a lack of transparency about the origins […]

18h ago
YouTube: Two Minute Papers

I Never Thought I’d See This Happen

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Navier-Stokes solution paper is available here: https://openai.com/index/navier-stokes-solution/ My fluid simulations and papers: https://users.cg.tuwien.ac.at/zsolnai/gfx/fluid_control_msc_thesis/ https://users.cg.tuwien.ac.at/zsolnai/gfx/real_time_fluid_control_eg/ The flow from simulation to reality: https://www.nature.com/articles/s41567-022-01788-5 All papers: https://users.cg.tuwien.ac.at/zsolnai/ F

20h ago
The Verge AI

Universal Music is launching an AI music platform with ElevenLabs

Universal Music Group is launching a new AI-powered platform that will allow users to draw from its catalog of licensed music to create song remixes, mashups, and new takes on tracks, according to an announcement on Thursday. The record label is developing the platform through a multiyear licensing agreement with ElevenLabs, a company that specializes […]

13h ago
MarkTechPost

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching service that sits between the application and […] The post Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster appeared first on MarkTechPost .

6h ago
The Verge AI

Slack can now vibe-code interactive charts and reports inside chats

A new feature coming to Slack will allow you to build interactive reports, polls, dashboards, presentations, microsites, and other tools directly inside a chat. With Slackforce Surfaces, you can describe to Slackbot what you need, and it will use AI to gather information from relevant conversations and connected apps, like Google Drive or Salesforce, to […]

7h ago
The Verge AI

Schools are catching on to Big Tech’s playbook

It's the hot new thing in tech, and it's where all the jobs are. Students who don't learn to use it fall behind. And to help them catch up in time, its creators are graciously providing the resources and curriculum for learning it, often pro bono. That's the narrative AI companies are pitching schools on […]

9h ago
The Verge AI

Why the current tech backlash feels different

This interview has been lightly edited for length and clarity.  Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. This is Nick Statt, senior producer. And I’m joined by our brand-new supervising producer, Greg Ott. Greg Ott: Good day, everyone. And Hi, Nilay. Nilay is here too. He is […]

15h ago
GitHub trending Python (daily) Free tool Open weight

NVIDIA/Megatron-LM

NVIDIA's Megatron-LM is trending on GitHub. It is an open research framework for training transformer models at scale across distributed hardware.

Hugging Face trending Free tool Open weight

nvidia/Qwen3.8-Flash-Next-NVFP4

NVIDIA released Qwen3.8-Flash-Next-NVFP4, an open image-text-to-text model optimized for inference on NVIDIA hardware. Free to download and use.

Hugging Face trending Free tool Open weight

MiniMaxAI/MiniMax-H3

MiniMax-H3, an open image-to-video model from MiniMaxAI, is trending on Hugging Face with over 5 million downloads. Free to download and run locally.

Hugging Face trending Free tool Open weight

Lightricks/LTX-2.5

Lightricks released LTX-2.5, an open-weight image-to-video model available on Hugging Face. Free to download and use.

GitHub trending (daily) Free tool Open weight

alsk1992/CloddsBot

CloddsBot is an open-source AI trading agent built on Claude that autonomously executes trades across 1,000+ markets including Polymarket, Kalshi, Binance, and Solana DEXs. Self-hosted and free to run.

GitHub trending Jupyter (daily) Free tool Open weight

FareedKhan-dev/all-agentic-architectures

all-agentic-architectures is a Python library bundling 35 production agentic patterns (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, and more) with multi-provider LLM support and a 17-task benchmark leaderboard. Useful reference for anyone building or evaluating agent systems.

GitHub trending Python (daily) Free tool Open weight

TauricResearch/TradingAgents

TradingAgents is an open multi-agent LLM framework for financial trading. It chains specialist agents for research, debate, and execution.

GitHub trending Jupyter (daily) Free tool Open weight

ageron/handson-ml2

Hands-On Machine Learning (2nd edition) is deprecated. Use handson-ml3 or handson-mlp instead if you rely on this textbook repo.

GitHub trending (daily) Free tool Open weight

ayghri/i-have-adhd

i-have-adhd is a skill for coding agents that formats output to be less overwhelming and easier to parse. Free to use.

Hugging Face trending Free tool Open weight

Viggle/Viggle-Animate

Viggle-Animate is a video-to-video model from Viggle released on Hugging Face. Free to download and use.

Hugging Face trending Free tool Open weight

m-a-p/YuE2-3B

YuE2-3B is a new open-weight text-to-audio model from m-a-p. Free to download and use.

Hugging Face daily papers Free tool

World in World: Explore the World with World Models

World in World is a new video world model that lets you explore a source video from new viewpoints while keeping the action synchronized. Enables flexible, long-horizon interactive exploration.

1d ago
Hugging Face daily papers Free tool

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

SchemeArena is a new benchmark for testing whether LLM agents will pursue hidden misaligned goals when given the chance. It factors in instrumental motivation, environment design, oversight, and perceived consequences to understand how scheming emerges.

3d ago
Hugging Face daily papers Free tool

Scaling Automatic Research Agents via World Models

Researchers scaled automatic research agents using world models, letting LLMs independently run experiments and learn from results. Post-training with reinforcement learning is the key to getting agents that can autonomously iterate on empirical problems.

Aug 29
Hugging Face daily papers Free tool

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-trivial challenge. This difficulty is largely attributed to two primary limitations: (1) the absence of structured relational modeling hinders agents from capturing complex, time-varying interactions among battlefield entities; and (2) conventional flat architectures often lack the capability to explicitly model tactical roles, lea

1d ago
Hugging Face daily papers Free tool

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization t

1d ago
Hugging Face daily papers Free tool

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduce PARSER, which decouples reading from reasoning. A bank of lightweight subagents each bound to a single chunk read the entire document in parallel, while a lead agent reasons in depth through iterative scatter--gather rounds: at each round it broadcasts a query to all subagents, aggregates the retu

5d ago
Hugging Face daily papers Free tool

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is stric

6d ago
Hugging Face daily papers Free tool

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a

2d ago
Hugging Face daily papers Free tool

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisso

3d ago
Hugging Face daily papers Free tool

TempCloze: Can Video-LLMs Identify the Missing Middle?

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Seman

Sep 1
Hugging Face daily papers Free tool

FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-wind

1d ago
Hugging Face daily papers Free tool

SenseNova-U1.5: Towards Native Unified Visual Intelligence

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their cap

1d ago
Hugging Face daily papers Free tool

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained s

2d ago
Hugging Face daily papers Free tool

Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digital ripple. We present Mi-Ripple, a diagnosis-guided restoration workflow that suppresses this digital ripple while protecting image structure. Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then combines selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration. This separation enables low-distortion filtering when artifacts are spectrally isolated and visual reconstruction when filtering would erase le

1d ago
Hugging Face daily papers Free tool

UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration

All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using a single universal model. Existing methods typically prioritize modeling inter-task heterogeneity (e.g., distinct data distributions and degradation types). However, they largely neglect the inherent homogeneity present in medical images, such as widely shared anatomical structures within and across modalities, which can be leveraged to ease model training and improve generalization. To this end, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and H

1d ago
Hugging Face daily papers Free tool

Generative Late-Interaction Embeddings For Visual Document Retrieval

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means ce

1d ago
Hugging Face daily papers Free tool

Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs

Code world models represent worlds as executable programs, but this representation alone does not determine how to construct a complex world. We introduce Recursive Code World Models (RCWM), a framework for reconstructing complex 3D worlds in code from a single reference image. RCWM couples a Recursive Scene Program (RSP) representation with a construction solver that recursively calls itself. An RSP represents the executable world as compositional scene code, while each solver call follows the same complete process: establish the whole, recursively reconstruct unresolved parts, and revisit th

1d ago
Hugging Face daily papers Free tool

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Con

2d ago
Hugging Face daily papers Free tool

The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements

We consider the problem of support recovery for sparse binary signals from noisy linear measurements. For sparse Gaussian measurement matrices we identify sufficient conditions on the minimal sample size for maximum-likelihood recovery in the high-SNR regime ds/p to infty, where p denotes the signal dimension, s the number of non-zero components of the signal, and d the expected number of non-zero components per row of measurement. Combined with known lower bounds, this yields an information-theoretic threshold of order slog(p/s) / log(ds/p), making explicit the price of measurement sparsity.

3d ago
Hugging Face daily papers Free tool

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive dev

4d ago
Hugging Face daily papers Free tool

CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation

Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermining clinician trust and clinical adoption readiness. We developed CARDEA, a unified large vision-language model that serves as the inference core of a CAG pipeline. It was trained solely on public datasets and closed-ended tasks in three stages: visual feature alignment, a self-distilled Chain-of-Bo

4d ago
Hugging Face daily papers Free tool

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, w

Sep 2
Hugging Face daily papers Free tool

HyQuant: Hybrid-Precision Quantization for LLM Attention

Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose HyQuant, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of ver

Aug 28

The Kitchen

What ran today and what it cost.

Today's spend $0.0600 12 calls · 230,104 tokens
Last 30 days $1.80 298 calls · 4,319,293 tokens
Full breakdown Open The Kitchen → Per-provider rollups, sparklines, model registry