LLM INFERENCE PERFORMANCE ANALYSIS ACROSS WINDOWS, WINDOWS SUBSYSTEM FOR LINUX 2 AND NATIVE LINUX

Authors

  • Muhammad Anshori Universitas Nurul Jadid Paiton, Probolinggo, Indonesia
  • Maulidiansyah Universitas Nurul Jadid Paiton
  • Moh Sukron Universitas Nurul Jadid Paiton

DOI:

https://doi.org/10.33480/jitk.v12i1.8435

Keywords:

Benchmarking, LLM Inference, Linux, Windows, WSL2

Abstract

The increasing use of Large Language Models (LLMs) on consumer devices raises practical concerns regarding inference latency and resource efficiency, which can vary depending on the operating system environment. However, it is still unknown which operating system is most optimal and efficient for executing local LLM workloads on hardware with limited resources. Therefore, this study aims to compare and evaluate the inference performance of LLMs across three operating system environments: native Windows, Windows Subsystem for Linux 2 (WSL2), and native Linux on a low-spec laptop equipped with a 12th Gen Intel Core i3 processor and 8 GB of RAM. The Ollama framework is used to run several lightweight LLM models, including Llama3.2:3B, Qwen2.5:3B, and Phi3.5:3.8B. An experimental benchmarking approach is conducted using three levels of prompt complexity (light, medium, and heavy), with 10 trials for each prompt scenario to ensure consistency of results. Evaluation metrics include inference latency, CPU utilization, and memory consumption during model execution. Results show that the operating system environment significantly impacts the inference performance of LLMs. For example, using the Llama3.2:3B model under heavy prompting, the average latency reaches 71.24 seconds on Windows, 75.96 seconds on WSL2, and 70.80 seconds on native Linux, while under light prompting the average latency is 11.76 seconds, 12.73 seconds, and 10.83 seconds, respectively. These findings indicate that native Linux generally provides more efficient LLM inference on resource-constrained hardware. This study contributes an empirical and repeatable comparison across operating system environments, providing practical insights for selecting deployment platforms for LLM applications.

Downloads

Download data is not yet available.

References

[1] K. Alizadeh et al., “LLM in a flash: Efficient Large Language Model Inference with Limited Memory,” Proc. Annu. Meet. Assoc. Comput. Linguist., vol. 1, pp. 12562–12584, 2024, doi: 10.18653/v1/2024.acl-long.678.

[2] W. X. Zhao, K. Zhou, and J. Li, “A Survey of Large Language Models,” Front. Comput. Sci., vol. 19, no. 7, pp. 1–144, 2025, doi: 10.1007/s11704-024-40663-9.

[3] J. Xiao, Q. Huang, X. Chen, and C. Tian, “Understanding Large Language Models in Your Pockets: Performance Study on COTS Mobile Devices,” vol. 14, no. 8, pp. 1–18, 2024, [Online]. Available: http://arxiv.org/abs/2410.03613

[4] Y. Shen, Y. Zhang, and D. Yuan, “A Hybrid Online and Offline Requests Inference Serving System for LLM in Private Computer Environment,” IEEE Trans. Comput., vol. 75, no. 3, pp. 888–900, 2025, doi: 10.1109/TC.2025.3643296.

[5] K. T. Chitty-Venkata et al., “LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators,” Proc. SC 2024-W Work. Int. Conf. High Perform. Comput. Networking, Storage Anal., pp. 1362–1379, 2024, doi: 10.1109/SCW63240.2024.00178.

[6] V. Rajesh, O. Jodhpurkar, P. Anbuselvan, M. Singh, and A. Jallepali, “Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS,” arXiv Prepr., no. Llm, 2026.

[7] H. Shen, H. Chang, B. Dong, Y. Luo, and H. Meng, “Efficient LLM Inference on CPUs,” no. NeurIPS, pp. 33–46, 2025, doi: 10.1007/978-3-031-85747-8_3.

[8] S. Samsi et al., “From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference,” 2023 IEEE High Perform. Extrem. Comput. Conf. HPEC 2023, 2023, doi: 10.1109/HPEC58863.2023.10363447.

[9] D. Huang and Z. Wang, “LLMs at the Edge: Performance and Efficiency Evaluation with Ollama on Diverse Hardware,” Proc. Int. Jt. Conf. Neural Networks, 2025, doi: 10.1109/IJCNN64981.2025.11228317.

[10] E. Chung, Y. Jia, A. Jezghani, and H. Kim, “Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference,” 2026, [Online]. Available: http://arxiv.org/abs/2603.22774

[11] Q. Song et al., A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources, vol. 1, no. 1. arXiv, 2026. [Online]. Available: http://arxiv.org/abs/2505.15030

[12] E. Dinçer and Z. H. Kilimci, “Real-Time and Offline Large Language Models on Edge Devices: A Systematic Review,” pp. 0–17, 2025, doi: 10.20944/preprints202512.2383.v1.

[13] D. Park and B. Egger, “Improving Throughput-Oriented LLM Inference with CPU Computations,” Parallel Archit. Compil. Tech. - Conf. Proceedings, PACT, no. March, pp. 233–245, 2024, doi: 10.1145/3656019.3676949.

[14] G. Qu, Q. Chen, W. Wei, Z. Lin, X. Chen, and K. Huang, “Mobile Edge Intelligence for Large Language Models: A Contemporary Survey,” IEEE Commun. Surv. Tutorials, vol. 27, no. 6, pp. 3820–3860, 2025, doi: 10.1109/COMST.2025.3527641.

[15] T. Nguyen and T. Nguyen, “An Evaluation of LLMs Inference on Popular Single-board Computers,” arXiv Prepr., 2025, [Online]. Available: http://arxiv.org/abs/2511.07425

[16] S. Na et al., “FlexInfer: Flexible LLM Inference with CPU Computations,” MLSys, 2025, [Online]. Available: https://mlsys.org/virtual/2025/poster/3234%0Ahttps://openreview.net/forum?id=sFNRNTduKO

[17] F. Liu, Z. Kang, and X. Han, “Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models,” Proc. 2024 3rd Int. Conf. Artif. Intell. Intell. Inf. Process. AIIIP 2024, no. March, pp. 152–159, 2025, doi: 10.1145/3707292.3707358.

[18] M. Lazuka, A. Anghel, and T. Parnell, “LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services,” Int. Conf. High Perform. Comput. Networking, Storage Anal. SC, 2024, doi: 10.1109/SC41406.2024.00022.

[19] H. Chen, C. Tian, Z. He, B. Yu, Y. Liu, and J. Cao, “Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric,” no. 2, 2025, [Online]. Available: http://arxiv.org/abs/2508.11269

[20] L. Stuhlmann, M. F. Argerich, and J. Fürst, “Bench360: Benchmarking Local LLM Inference from 360 Degrees,” 2026, [Online]. Available: http://arxiv.org/abs/2511.16682

[21] Z. Yuan et al., “LLM Inference Unveiled: Survey and Roofline Model Insights,” 2024, [Online]. Available: http://arxiv.org/abs/2402.16363

[22] W. Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” SOSP 2023 - Proc. 29th ACM Symp. Oper. Syst. Princ., pp. 611–626, 2023, doi: 10.1145/3600006.3613165

Downloads

Published

2026-08-31

How to Cite

[1]
“LLM INFERENCE PERFORMANCE ANALYSIS ACROSS WINDOWS, WINDOWS SUBSYSTEM FOR LINUX 2 AND NATIVE LINUX”, jitk, vol. 12, no. 1, pp. 451–462, Aug. 2026, doi: 10.33480/jitk.v12i1.8435.

Most read articles by the same author(s)