Instruction Learning Paradigms: A Dual Perspective on White
Abstract
Optimizing instructions for large language models (LLMs) is critical for harnessing their full potential in complex and diverse tasks. However, relying solely on white-box approaches demands extensive computational resources and offers limited representational capacity, while black-box models can incur prohibitive financial costs. To address these challenges, we introduce a novel framework that seamlessly merges the strengths of both paradigms. Black-box models provide high-quality, diverse instruction initializations, and white-box models supply fine-grained interpretability through hidden states and output features. By enforcing a semantic similarity constraint, these components fuse into a unified high-dimensional representation that captures deep semantic and structural nuances, enabling an iterative optimization process to refine instruction quality and adaptability. The framework aligns black-box and white-box representations within a shared semantic space and incorporates a feature adaptation module to dynamically capture inter-model consistency and complementarity, thereby enhancing information fusion and semantic robustness. Extensive evaluations across a broad spectrum of benchmark tasks—ranging from complex reasoning to three English-to-X translation tasks—demonstrate that our approach consistently outperforms strong baselines on average. This fusion of black-box initialization with semantic refinement yields an effective and scalable solution for instruction optimization across diverse benchmark tasks.
Similar content being viewed by others
Data Availability
No datasets were generated or analysed during the current study.
References
Bai, G., Chai, Z., Ling, C., Wang, S., Lu, J., Zhang, N., Shi, T., Yu, Z., Zhu, M., Zhang, Y., Song, X., Yang, C., Cheng, Y., & Zhao, L. (2024a). Beyond efficiency: A systematic survey of resource-efficient large language models. arXiv preprint arXiv:2401.00625 https://doi.org/10.48550/arXiv.2401.00625
Bai, Y., Du, X., Liang, Y., Jin, Y., Zhou, J., Liu, Z., Fang, F., Chang, M., Zheng, T., Zhang, X., Ma, N., Wang, Z., Yuan, R., Wu, H., Lin, H., Huang, W., Zhang, J., Lin, C., Fu, J., Zhang, G. (2024b). COIG-CQIA: Quality is all you need for Chinese instruction fine-tuning. arXiv preprint arXiv:2403.18058 https://doi.org/10.48550/arXiv.2403.18058
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://doi.org/10.48550/arXiv.2005.14165
Cai, W., Jiang, J., Wang, F., Tang, J., Kim, S., & Huang, J. (2025). A survey on mixture of experts in large language models. IEEE Transactions on Knowledge and Data Engineering, 37(7), 3896–3915. https://doi.org/10.1109/TKDE.2025.3554028
Cao, Y., Kang, Y., Wang, C., & Sun, L. (2023). Instruction mining: Instruction data selection for tuning large language models. arXiv preprint arXiv:2307.06290 https://doi.org/10.48550/arXiv.2307.06290
Chen, L., Chen, J., Goldstein, T., Huang, H., & Zhou, T. (2023). Instructzero: Efficient instruction optimization for black-box large language models. arXiv preprint arXiv:2306.03082.
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., & Xing, E. P. (2023, March 30). Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality. LMSYS Org https://www.lmsys.org/blog/2023-03-30-vicuna/
Chirkova, N., & Nikoulina, V. (2024). Zero-shot cross-lingual transfer in instruction tuning of large language models arXiv preprint. arXiv:2402.14778.
Conti, E., Madhavan, V., Petroski Such, F., Lehman, J., Stanley, K., & Clune, J. (2018). Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. Advances in Neural Information Processing Systems, 31 https://proceedings.neurips.cc/paper/2018/hash/b1301141feffabac455e1f90a7de2054-Abstract.html
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P. N., & Hoi, S. (2023). Instructblip: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems, 36, 49250–49267.
Feng, S., Fang, G., Ma, X., & Wang, X. (2025). Efficient reasoning models: A survey. arXiv preprint arXiv:2504.10903.
Ge, Y., Hua, W., Mei, K., Ji, J., Tan, J., Xu, S., Li, Z., & Zhang, Y. (2023). Openagi: When llm meets domain experts. Advances in Neural Information Processing Systems, 36, 5539–5568
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., & Yang, Y. (2023). Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. arXiv preprint arXiv:2309.08532.
Hake, J., Crowley, M., Coy, A., Shanks, D., Eoff, A., Kirmer-Voss, K., Dhanda, G., & Parente, D. J. (2024). Quality, accuracy, and bias in chatgpt-based summarization of medical abstracts. The Annals of Family Medicine, 22(2), 113–120.
Heo, J., Xiong, M., Heinze-Deml, C., & Narain, J. (2024). Do llms estimate uncertainty well in instruction-following? arXiv preprint arXiv:2410.14582.
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., … Sifre, L. (2022). An empirical analysis of compute-optimal large language model training. Advances in Neural Information Processing Systems, 35, 30016–30030 https://doi.org/10.52202/068431-2176
Huang, S., Yang, K., Qi, S., & Wang, R. (2024). When large language model meets optimization. arXiv preprint arXiv:2405.10098.
Hu, L., Liu, Z., Zhao, Z., Hou, L., Nie, L., & Li, J. (2024). A survey of knowledge enhanced pre-trained language models. IEEE Transactions on Knowledge and Data Engineering, 36(4), 1413–1430. https://doi.org/10.1109/TKDE.2023.3310002
Hu, Z., Zhang, Y., Xiao, M., Wang, W., Feng, F., & He, X. (2025). Exact and efficient unlearning for large language model-based recommendation. IEEE Transactions on Knowledge and Data Engineering, 37(10), 5866–5877. https://doi.org/10.1109/TKDE.2025.3594687
Jin, B., Liu, G., Han, C., Jiang, M., Ji, H., & Han, J. (2024). Large language models on graphs: A comprehensive survey. IEEE Transactions on Knowledge and Data Engineering, 36(12), 8622–8642. https://doi.org/10.1109/TKDE.2024.3469578
Kung, P.-N., & Peng, N. (2023). Do models really learn to follow instructions? an empirical study of instruction tuning. arXiv preprint arXiv:2305.11383.
Li, J., Liu, W., Ding, Z., Fan, W., Li, Y., & Li, Q. (2025). Large language models are in-context molecule learners. IEEE Transactions on Knowledge and Data Engineering, 37(7), 4131–4143. https://doi.org/10.1109/TKDE.2025.3557697
Li, Z., Peng, B., He, P., & Yan, X. (2023). Evaluating the instruction-following robustness of large language models to prompt injection. arXiv preprint arXiv:2308.10819.
Lin, X., Wu, Z., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., & Low, B. K. H. (2024). Use your instinct: Instruction optimization for llms using neural bandits coupled with transformers. Forty-first International Conference on Machine Learning.
Ling, C., Zhao, X., Lu, J., Deng, C., Zheng, C., Wang, J., Chowdhury, T., Li, Y., Cui, H., Zhang, X., Zhao, T., Panalkar, A., Mehta, D., Pasquali, S., Cheng, W., Wang, H., Liu, Y., Chen, Z., Chen, H., Zhao, L. (2025). Domain specialization as the key to make large language models disruptive: A comprehensive survey. ACM Computing Surveys, 58(3), 1–39 https://doi.org/10.1145/3764579
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., & Ruan, C. (2024b). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.
Liu, B., Chen, C., Gong, Z., Liao, C., Wang, H., Lei, Z., Liang, M., Chen, D., Shen, M., Zhou, H., Jiang, W., Yu, H., & Li, J. (2024). MFTCoder: Boosting code LLMs with multitask fine-tuning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5430–5441). Association for Computing Machinery https://doi.org/10.1145/3637528.3671609
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., & Roberts, A. (2023). The Flan collection: Designing data and methods for effective instruction tuning. Proceedings of the 40th International Conference on Machine Learning, 202, 22631–22648 https://proceedings.mlr.press/v202/longpre23a.html
Lou, R., Zhang, K., & Yin, W. (2024). Large language model instruction following: A survey of progresses and challenges. Computational Linguistics, 50(3), 1053–1095 https://doi.org/10.1162/coli_a_00523
Muhtar, D., Li, Z., Gu, F., Zhang, X., & Xiao, P. (2024). Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model. European Conference on Computer Vision (pp. 440–457). Springer.
Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., & Mian, A. (2023). A comprehensive overview of large language models. arXiv preprint. arXiv:2307.06435
OpenAI: ChatGPT. May 24 version [Large language model] (2023). https://chat.openai.com
Pan, H. (2025). Information Extraction From Scientific Literature. Temple University.
Song, S., Li, X., Li, S., Zhao, S., Yu, J., Ma, J., Mao, X., Zhang, W., & Wang, M. (2025). How to bridge the gap between modalities: Survey on multimodal large language model. IEEE Transactions on Knowledge and Data Engineering, 37(9), 5311–5329. https://doi.org/10.1109/TKDE.2025.3527978
Sun, X., Dong, L., Li, X., Wan, Z., Wang, S., Zhang, T., Li, J., Cheng, F., Lyu, L., Wu, F., & Wang, G. (2023). Pushing the limits of chatgpt on nlp tasks. arXiv preprint arXiv:2306.09719.
Sun, X., Shi, K., Tang, H., Wang, D., Xu, G., & Li, Q. (2025). Educating language models as promoters: Multi-aspect instruction alignment with self-augmentation. IEEE Transactions on Knowledge and Data Engineering, 37(8), 4564–4577. https://doi.org/10.1109/TKDE.2025.3569585
Su, T., Zhang, J., Yu, Z., Wang, G., & Liu, X. (2023). Stkd: Distilling knowledge from synchronous teaching for efficient model compression. IEEE Transactions on Neural Networks and Learning Systems, 34(12), 10051–10064. https://doi.org/10.1109/TNNLS.2022.3164264
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., & Hashimoto, T. B. (2023, March 13). Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models https://crfm.stanford.edu/2023/03/13/alpaca
Wang, H., Li, W., Xia, X.-G., & Du, Q. (2025). Bihot: A large-scale dataset and benchmark for hyperspectral camouflaged object tracking. IEEE Transactions on Neural Networks and Learning Systems, 36(9), 16392–16406. https://doi.org/10.1109/TNNLS.2025.3564059
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H. G., Purohit, I., Mondal, I., Anderson, J., Kuznia, K., Doshi, K., Patel, M., Pal, K. K., Moradshahi, M., Parmar, M., Purohit, M., Varshney, N., Kaza, P. R., Verma, P., Puri, R. S., Karia, R., Sampat, S. K.,Doshi, S., Mishra, S., Reddy, S., Patro, S., Dixit, T., Shen, X., Baral, C., Choi, Y.,Smith, N. A., Hajishirzi, H., & Khashabi, D. (2022). Super-naturalinstructions: Generalization via declarative instructions. on 1600+ nlp tasks. arXiv preprint http://arxiv.org/abs/2204.07705.
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., & Le, Q. V. (2021). Finetuned language models are zero-shot learners arXiv preprint. arXiv:2109.01652.
Wen, B., Ke, P., Gu, X., Wu, L., Huang, H., Zhou, J., Li, W., Hu, B., Gao, W., Xu, J., Liu, Y., Tang, J., Wang, H., & Huang, M. (2024). Benchmarking complex instruction-following with multiple constraints composition. Advances in Neural Information Processing Systems, 37, 137610–137645
White, J., Hays, S., Fu, Q., Spencer-Smith, J., Schmidt, D.C. (2024). ChatGPT Prompt Patterns for Improving Code Quality, Refactoring, Requirements Elicitation, and Software Design. In: Nguyen-Duc, A., Abrahamsson, P., Khomh, F. (eds) Generative AI for Effective Software Development. Springer, Cham. https://doi.org/10.1007/978-3-031-55642-5_4
Wu, X., Wang, M., Liu, Y., Shi, X., Yan, H., Lu, X., Zhu, J., & Zhang, W. (2024b). Lifbench: Evaluating the instruction following performance and stability of large language models in long-context scenarios https://doi.org/10.48550/arXiv.2411.07037 arXiv preprint.
Wu, Z., Dadu, A., Nalls, M., Faghri, F., & Sun, J. (2024a). Instruction tuning large language models to understand electronic health records. Advances in Neural Information Processing Systems, 37, 54772–54786.
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., & Jiang, D. (2023). Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint http://arxiv.org/abs/2304.12244.
Xu, F. F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z. Z., Zhou, X., Guo, Z., & Cao, M. (2024). Theagentcompany: benchmarking llm agents on consequential real world tasks. arXiv preprint http://arxiv.org/abs/2412.14161.
Yu, Z., Zhang, X., Shang, N., Huang, Y., Xu, C., Zhao, Y., Hu, W., & Yin, Q. (2023). Wavecoder: Widespread and versatile enhanced instruction tuning with refined data generation. arXiv preprint. http://arxiv.org/abs/2312.14187
Zaleppa, P., Kaza, S., & Taylor, B. (2024). Generating an instruction dataset to build cyber intelligent large language models. In: 2024 Cyber Awareness and Research Symposium (CARS), pp. 1–5. https://doi.org/10.1109/CARS61786.2024.10778717.
Zhang, J., Vahidian, S., Kuo, M., Li, C., Zhang, R., Yu, T., Wang, G., & Chen, Y. (2024a). Towards building the federatedgpt: Federated instruction tuning. In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6915–6919 IEEE.
Zhang, Y.-F., Zhang, H., Tian, H., Fu, C., Zhang, S., Wu, J., Li, F., Wang, K., Wen, Q., Zhang, Z., Wang, L., Jin, R., & Tan, T. (2024b). Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans? arXiv preprint http://arxiv.org/abs/2408.13257.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., … Wen, J.-R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223 https://doi.org/10.48550/arXiv.2303.18223
Zhao, Z., Wallace, E., Feng, S., Klein, D., & Singh, S. (2021). Calibrate before use: Improving few-shot performance of language models. International Conference on Machine Learning (pp. 12697–12706) PMLR.
Zheng, Y., Chen, Y., Qian, B., Shi, X., Shu, Y., & Chen, J. (2025). A review on edge large language models: Design, execution, and applications. ACM Computing Surveys, 57(8), 1–35.
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., & Levy, O. (2023). LIMA: Less is more for alignment. Advances in Neural Information Processing Systems, 36, 55006–55021 https://doi.org/10.52202/075280-2400
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., & Ba, J. (2022). Large language models are human-level prompt engineers. arXiv preprint http://arxiv.org/abs/2211.01910.
Acknowledgments
This work was supported by National Natural Science Foundation of China:(No.62441617), National Key R&D Program of China under the Strategic Science and Technology Innovation Cooperation Special Project (No. 2025YFE0209100), and CCF-Kuaishou Fund for Exploring Large Models (No.CCFKuaiShou2025012).
Author information
Authors and Affiliations
Contributions
Yanwei Ren wrote the main manuscript text, implemented the method, performed the experiments, and prepared all figures. Liu Liu supervised the project and revised the manuscript.All authors reviewed the manuscript.
Corresponding author
Additional information
Editor: Marco Lippi.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Appendices
Appendix A: Additional Statistical Results
To provide a clearer view of seed-level variability, Table 7 reports the standard deviations of accuracy across three different random seeds for all compared methods on the 30-task benchmark. The corresponding mean accuracies are reported in the main results table. Lower values indicate more stable performance across different runs.
Appendix B: Initialization Prompt
To guide ChatGPT in generating high-quality instructions, we utilized a structured initialization prompt. This prompt instructs the model to formulate 40 distinct instructions, each beginning with the phrase “The instruction was to,” while varying in style, structure, and complexity. The instructions must serve as high-level guiding principles without explicitly referencing input–output examples. Instead, they should be diverse in phrasing and sentence construction to ensure robustness and generalizability.
The Fig. 8 is the complete initialization prompt along with representative {init_token} examples from the taxonomy_animal task.
Appendix C: Best Instruction Across 30 Tasks
The Table 8 provides a comprehensive summary of the best-performing instructions identified for each of the 30 tasks in our evaluation. These instructions were discovered through an iterative optimization process, meaning they represent the prompt formulations that achieved the highest performance for their respective tasks. In each case, the instruction designated as “best” is the result of refining the prompt wording to maximally align the model’s output with the task’s objective. The tasks themselves span a broad range—from linguistic transformations and classification challenges to code execution and arithmetic operations—underscoring the versatility of our approach. The original results (summarized in Table VI) enumerate each task alongside its optimized instruction, illustrating how diverse tasks often require distinct prompt strategies to reach optimal performance.
The optimized instructions exhibit clear adaptation to the nature of each task through their wording and focus. For instance, tasks involving text transformation (such as converting an active-voice sentence to passive voice or informal language to formal) had instructions that explicitly stated the required transformation. The best instruction for the active-to-passive task directed the model to “rewrite the sentence to shift attention from the initiator of the action to the entity upon which the action is performed”, clearly describing the intended rephrasing. Similarly, the optimal prompt for a debugging task instructed the model to “use the Python interpreter to execute each line of code and print the output”, effectively guiding it through a step-by-step code execution procedure. These examples show how the highest-performing instructions encapsulate the exact operation or criterion needed for each task, whether it involves rephrasing text, performing a calculation, or simulating a process. By aligning the prompt’s content and style with the specific demands of the task, the instructions help focus the model’s reasoning on relevant actions and information, thereby improving task performance.
Across all tasks, a unifying characteristic of the optimized prompts is their clarity and specificity. Nearly all of the top instructions begin with a strong imperative verb (e.g., “Rewrite,” “Identify,” “Convert”) that immediately signals the required action to the model. This direct style ensures that the model’s output is closely aligned with the task goal, as the prompt precisely delineates what operation to perform. Furthermore, the instructions are concise yet detailed enough to avoid ambiguity. They often include explicit references to the key content or criteria of the task – for example, explicitly mentioning a “common category” to identify, the “first letter” to extract, or a “single-letter hint” to use in solving a puzzle. By highlighting such crucial details within the instruction, the prompts leave little room for misinterpretation, guiding the model along the intended solution path. This emphasis on unambiguous, task-focused language is consistent with best practices in prompt design and likely contributes substantially to their effectiveness.
Although all optimized instructions maintain clarity, they are not uniform in format; instead, their structural styles vary notably across different tasks. Some tasks are best addressed with a brief, direct command, whereas others benefit from a more elaborate, step-wise directive. For example, certain optimal prompts resemble a short procedure or pseudo-code, especially for tasks involving code execution or multi-step reasoning – a cue that a procedural tone can help the model follow the logical steps required. In contrast, simpler classification or conversion tasks often required nothing more than a single succinct sentence in the imperative mood to achieve top performance. This structural diversity suggests that the form of an effective instruction is highly task-dependent, with the prompt’s style being tailored to the nature of the problem. Importantly, however, even the more complex or structured instructions remain focused and unambiguous. In every case, the optimized prompt avoids extraneous information and zeroes in on the task, indicating that adapting the instruction’s form to the task can be done without sacrificing clarity. Such findings reinforce the idea that flexible prompt formulation – adjusting tone, detail, or format to fit the task – is key to eliciting the best possible performance from the model.
Several notable trends emerge when examining how these optimized instructions align with improved model performance. One clear pattern is the inclusion of domain-specific or task-specific cues directly in the instruction, which helps anchor the model’s reasoning in the correct context. For example, the best prompt for a chemical knowledge task explicitly mentions the element’s “location in the periodic chart,” immediately pointing the model to consider periodic table information. Similarly, instructions for tasks dealing with letter positions or word puzzles explicitly reference those details: a prompt might instruct “output the first letter of each word” or “provide a spaced sequence of characters” for reconstruction, ensuring the model knows to focus on letter-level operations. Another tactic observed is framing certain tasks as if they were programming or tool-use tasks. For instance, the optimized prompts for generating plural forms or synonyms asked the model to “create a program that takes a word as input and outputs” the transformed word. This approach effectively guides the model to produce a structured, single-word answer (as a program would), which is well-suited for tasks requiring a specific output format. These patterns indicate that performance gains were often achieved by making the instructions echo the structure of the solution or the process needed to obtain it. In other words, the prompt itself acts as a scaffold for the task: by leveraging the model’s strengths in following explicit procedural or rule-based instructions, the optimized prompts channel the model’s behavior towards the correct solution strategy. Such insights underscore how nuanced prompt formulation – down to including relevant context and mimicking procedural frameworks – can strongly influence the quality of an LLM’s output.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Ren, Y., Liu, L., Yu, B. et al. Instruction Learning Paradigms: A Dual Perspective on White-Box and Black-Box LLMs. Mach Learn 115, 200 (2026). https://doi.org/10.1007/s10994-026-07137-4
Received:
Revised:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s10994-026-07137-4
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.