DOI:
https://doi.org/10.14483/23448350.24600Published:
07/23/2026Issue:
Vol. 53 No. 1 (2026): Vol. 53 No. 1(2026): January-April 2026Section:
Research ArticlesAnalysis of Vector Chunk and Overlap Sizes in Optimizing Vector Databases in Artificial Intelligence Agent Automation
Análisis del Tamaño de Fragmentos y Solapamiento de Vectores en la Optimización de Bases de Datos Vectoriales para la Automatización de Agentes de IA
Keywords:
Artificial Intelligence, Natural Language Processing, Retrieval-Augmented Generation, Large Language Models, Information Retrieval, Document Vectorization, Unstructured Data, AI Agent Automation (en).Keywords:
Inteligencia Artificial, Procesamiento del Lenguaje Natural, Generación Aumentada por Recuperación (RAG), Modelos de Lenguaje a Gran Escala (LLMs), Recuperación de Información, Vectorización de Documentos, Datos No Estructurados, Automatización de Agentes de IA (es).Downloads
Abstract (en)
Recent advances in language models and retrieval-augmented generation (RAG) systems have highlighted the importance of optimizing parameters such as chunk size and overlap when vectorizing unstructured data for efficient information retrieval. This study explores the impact of varying chunking configurations on the performance of a RAG system designed to answer queries regarding the Regulations of the Social Service of the Faculty of Science, UNAM. Using N8N as the integration platform and OPENAI gpt-4o-mini as the LLM tool, different chunking strategies were evaluated based on four key metrics: correctness, semantic similarity, context relevance, and answer relevance. Results suggest that while overall performance differences between chunking strategies were not statistically significant, variations in chunk size and overlap can influence the consistency and quality of responses, particularly regarding answer relevance. These findings offer practical guidelines for optimizing chunking parameters in a RAG system for retrieving information from regulations documents to enhance retrieval accuracy and response reliability.
Abstract (es)
Avances recientes en modelos de lenguaje y sistemas de generación aumentada por recuperación (RAG) han destacado la importancia de optimizar parámetros como el tamaño de los fragmentos (chunk size) y su solapamiento (overlap) al vectorizar datos no estructurados para una recuperación de información eficiente. Este estudio explora el impacto de diferentes configuraciones de fragmentación en el rendimiento de un sistema RAG diseñado para responder consultas sobre el Reglamento del Servicio Social de la Facultad de Ciencias de la UNAM. Utilizando N8N como plataforma de integración y OPENAI gpt-4o-mini como herramienta de LLM, se evaluaron diferentes estrategias de fragmentación con base en cuatro métricas clave: corrección, similitud semántica, relevancia del contexto y relevancia de la respuesta. Los resultados sugieren que, si bien las diferencias generales de rendimiento entre las estrategias de fragmentación no fueron estadísticamente significativas, las variaciones en el tamaño de los fragmentos y su solapamiento pueden influir en la consistencia y calidad de las respuestas, particularmente en lo que respecta a la relevancia de la respuesta. Estos hallazgos ofrecen lineamientos prácticos para optimizar los parámetros de fragmentación en un sistema RAG destinado a recuperar información de documentos reglamentarios, con el fin de mejorar la precisión de la recuperación y la confiabilidad de las respuestas.
References
Chen, T., Wang, H., Chen, S., Yu, W., Ma, K., Zhao, X., Zhang, H., & Yu, D. (2024). Dense X retrieval: What retrieval granularity should we use? arXiv preprint. https://doi.org/10.48550/arXiv.2312.06648
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint. https://doi.org/10.48550/arXiv.1810.04805
Duarte, A. V., Marques, J., Graça, M., Freire, M., Li, L., & Oliveira, A. L. (2024). LumberChunker: Long-form narrative document segmentation. arXiv preprint. https://doi.org/10.48550/arXiv.2406.17526
Google Cloud. (n.d.). What are AI agents? Definition, examples, and types. https://cloud.google.com/discover/what-are-ai-agents
Juvekar, K., & Purwar, A. (2024). Introducing a new hyper-parameter for RAG: Context window utilization. arXiv preprint. https://doi.org/10.48550/arXiv.2407.19794
Keselman, H. J., Algina, J., & Kowalchuk, R. K. (2001). The analysis of repeated measures designs: A review. British Journal of Mathematical and Statistical Psychology, 54(1), 1-20. https://doi.org/10.1348/000711001159357
Kimothi, A. (2024, September 23). Breaking it down: Chunking for better RAG. Medium.
https://medium.com/data-science/breaking-it-down-chunking-techniques-for-better-rag-3fd288bf25a0
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv preprint. https://doi.org/10.48550/arXiv.2005.11401
Marcus, G. (2020). The next decade in AI: The steps towards robust artificial intelligence. arXiv preprint.
https://doi.org/10.48550/arXiv.2002.06177
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint. https://doi.org/10.48550/arXiv.1301.3781
OpenAI (n.d.). What are tokens and how to count them? OpenAI Help Center.
https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
Pinecone (n.d.). Vector similarity explained. https://www.pinecone.io/learn/vector-similarity/
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI. https://openai.com/research/language-unsupervised
Sheldon, R., & Yasar, K. (2024, November 13). What is an AI prompt? TechTarget.
https://www.techtarget.com/searchenterpriseai/definition/AI-prompt
Singh, I. S., Aggarwal, R., Allahverdiyev, I., Taha, M., Akalin, A., Zhu, K., & O'Brien, S. (2024). ChunkRAG: Novel LLM-chunk filtering method for RAG systems. arXiv. https://doi.org/10.48550/arXiv.2410.19572
Stryker, C., & Holdsworth, J. (2024, August 11). What is NLP (natural language processing)? IBM Think.
https://www.ibm.com/think/topics/natural-language-processing
Superteams.ai (2025). A deep dive into chunking strategy, chunking methods, and precision in RAG applications.
Taipalus, T. (2024). Vector database management systems: Fundamental concepts, use-cases, and current challenges. Cognitive Systems Research, 85, 101216. https://doi.org/10.1016/j.cogsys.2024.101216
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint. https://doi.org/10.48550/arXiv.2302.13971
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., Wang, L., Luu, A. T., Bi, W., Shi, F., & Shi, S. (2023). Siren's song in the AI ocean: A survey on hallucination in large language models. arXiv preprint. https://doi.org/10.48550/arXiv.2309.01219
Zhong, Z., Liu, H., Cui, X., Zhang, X., & Qin, Z. (2024). Mix-of-granularity: Optimize the chunking granularity for retrieval-augmented generation. arXiv preprint. https://doi.org/10.48550/arXiv.2406.00456
Zhou, J., Zhang, G., Alfarraj, O., Li, X., Tolba, A., & Zhang, H. (2024). DC-Graph: A chunk optimization model based on document classification and graph learning. Artificial Intelligence Review, 57, Article 143.
How to Cite
APA
ACM
ACS
ABNT
Chicago
Harvard
IEEE
MLA
Turabian
Vancouver
Download Citation
License
Copyright (c) 2026 Francisco Valdés-Souto, Daniela Richard-Ramírez, Emilio Domínguez-Valenzuela, Manuel Rodríguez-Flores, Natalia Edith Mejía-Bautista

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
When submitting their article to the Scientific Journal, the author(s) certifies that their manuscript has not been, nor will it be, presented or published in any other scientific journal.
Within the editorial policies established for the Scientific Journal, costs are not established at any stage of the editorial process, the submission of articles, the editing, publication and subsequent downloading of the contents is free of charge, since the journal is a non-profit academic publication. profit.