BACKGROUND: Large language models (LLMs) are increasingly being used to support diagnostic decision-making. It is unclear, however, whether their performance on knowledge and examination tasks translates to interpreting clinical case vignettes, especially when a complete diagnosis requires specifying several related elements, such as triggers, organ manifestations, or complications. METHODS: to assess six current language models from the Anthropic, Google, and Open AI companies. The strongest model of each company as of March 2026 was tested, as well as one smaller model from each. The models were required to provide a differential diagnosis with five disease entities and to name the most likely diagnosis. The answers were evaluated using a three-level scale. We drew a distinction between diagnoses with a single component and diagnoses with multiple components. RESULTS: The most likely diagnosis was most often correctly named by Claude Opus 4.6 (86.1%, 95% confidence interval [78.1; 91.6]), with GPT-5.4 close behind (85.1% [76.9; 90.8]). When partly correct answers were also counted, the three best-performing models all scored in the narrow range of 91.1-92.1%. The answers were highly stable. For diagnoses with multiple components, incorrect answers were more often incomplete than entirely wrong. CONCLUSION: LLMs have high diagnostic accuracy when challenged with published nephrological case scenarios. The interpretation of these findings is limited, however, because previous exposure of the models to individual vignettes, or parts thereof, cannot be ruled out. For clinical use, it must be borne in mind that plausible answers can be incomplete. Global accuracy measures may overestimate the diagnostic reliability of AI language models. The completeness and causal plausibility of the answers and their relevance to clinical action should be systematically investigated in benchmark studies.