25 papers this week in AI for science & health12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

AI for science & health research

Latest Paper· AI for science & health

Ocean emulator runs on unstructured mesh, beating baselines for currents

The authors present HClimRep-Ocean, a machine-learning emulator that operates directly on the unstructured computational mesh of the FESOM2 ocean model. Trained on a 209-year control simulation, the model forecasts ocean currents more accurately than all reference methods at 30-day lead times, though temperature and salinity predictions are less skilled than a simple damped-anomaly persistence forecast. The authors also demonstrate competitive performance on the OceanBench benchmark against a reanalysis product.

Kacper Nowak, Aleksei Koldunov, Nikolay Koldunov +7
AI for science & health

Joint deterministic-generative model extends reliable extreme-precipitation nowcasting to six hours

The authors present MW-Nowcast, a six-hour ensemble radar nowcasting model that jointly learns a deterministic predictor for organized precipitation structure and a generative flow-matching model for local residuals. The model doubles the available warning time for the most intense rainfall, delivering 6 h forecasts with skill previously confined to 3 h for the leading generative baseline.

today
Evaluation & benchmarks

LLMs struggle to match multidisciplinary tumor board discussions in new benchmark

The authors introduce OpenTumorBoard, a benchmark derived from 12,534 minutes of publicly available tumor board recordings on YouTube, comprising 611 patient cases and 19,157 discussion turns across ten specialist roles. Evaluating 14 general-purpose and medical LLMs, they find that the best models achieve modest scores of 3.43/5 on responding to individual specialist questions and 2.78/5 on generating entire board discussions that align with recorded consensus. Supervised finetuning and reinforcement learning on the benchmark data improve performance on a held-out test set.

today
Evaluation & benchmarks

Clinical language models diagnose well but often mishandle full encounters

Fang et al. introduce a clinician-authored benchmark that evaluates the process of clinical assessment rather than diagnosis from a complete case summary. Across 31 models, diagnostic accuracy reached 90.7%, yet even the strongest systems passed fewer than 30% of tasks when required to question a virtual patient, request appropriate examinations, respect constraints and provide the correct diagnosis.

2 days ago

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses