Audio-Visual Chilean Speaker Dataset
Mi proyecto de investigación aborda la ausencia de recursos de datos audiovisuales orientados a variantes lingüísticas locales, centrándose específicamente en el español hablado en Chile. A diferencia de grandes datasets internacionales (como VoxCeleb) enfocados mayoritariamente en inglés o chino sin distinción de acentos regionales, este trabajo implementa un pipeline automatizado de extremo a extremo utilizando recursos de fácil acceso y hardware de consumidor.
La innovación central radica en la recolección automatizada de Personas de Interés (POI) mediante un web scraper de noticias nacionales y procesamiento de lenguaje natural (PLN) para la extracción de entidades nombradas. El sistema complementa esta fase con la descarga masiva de material audiovisual desde plataformas de streaming, aplicando técnicas avanzadas de visión por computadora para la detección de rostros, seguimiento temporal (tracking), detección de transiciones de plano y optimización de parámetros espaciotemporales.
- Desarrollo de Pipeline Semiautomático: Creación de una arquitectura modular de código abierto capaz de automatizar la recopilación, filtrado, normalización y procesamiento de datos audiovisuales a partir de fuentes web abiertas.
- Extracción Masiva de Corpus Textual y POIs: Implementación de scrapers que recolectaron más de 420,000 artículos de noticias de cinco medios nacionales, detectando más de 420,000 entidades nombradas (PER) y validando perfiles de hablantes chilenos mediante ranking TF-IDF y bases de conocimiento.
- Recopilación Audiovisual a Gran Escala: Descarga exitosa de 8,303 archivos de video que totalizan 1.8 TB de almacenamiento y más de 121 días de metraje continuo en condiciones reales (in-the-wild).
- Optimización de Parámetros de Procesamiento: Determinación experimental de umbrales clave, demostrando que el salto de fotogramas a 15 FPS minimiza el error cuadrático medio (MSE), mientras que el escalado a 550 px acelera el procesamiento preservando el 86% de las caras detectadas.
- Impulso a la Colaboración Interdisciplinaria: Apertura de los datos y herramientas para fomentar la investigación en áreas de informática (generación de avatares, sincronización de labios), lingüística (estudio de la lengua hablada fuera del laboratorio) y ciencias sociales.
with speech fragments of Chilean Speakeres taken from YouTube, a special focus on fragments of News and other TV videos.
- A text corpus of Chilean News, first part uses directly scrapped News articles from Chilean sites. It will be cleaned and standardized, it will have monthly-packaged news bundles and the news will also be made available as a SaaS and/or REST API, with at least 48h desyncronization with the source site. No news are my property, i only provide the cleanup and tabulation service.
- Photogrammetry Toy: After having some problems with Meshroom, I am making my own Photogrammetry tool for fun. From a collection of photos to point clouds. Some nice handle of recomputing and warm-start of camera positions.
- Tarot RAG: Tarot cards are archetipal and have a lot of related concepts. This is a toy app that uses LLMs to model each tarot card as a document embedding. The idea is that you draw the cards and they can be understood as a sampling of documents, with an LLM returning an explanation and user-readable explanation of the docs. The other leg of the project is a Tarot translator, the idea is the reverse problem: You have a statement and try to express the same concepts using a tarot card draw. Similar to an emoji translator.
- SVG RL app/saas: New project just in the planning phase. i had to make a school site. I am no visual designer, so gen ai was useful to make the drawings..but they were raster and i needed svgs! So I manually traced the drawings and it was time consuming. So! Why didnt I use the auto-svg from inkscape? its awfull. I dont want all borders to be paths! Why didn’t you just use the borders! So RL to make it…
Agrwgar keyword ETL al curriculum y la página
Working in these project right now! Proyectos actuales: