AI Researcher and Engineer
PhD Researcher in Data Science and AI, University of Hull
Agentic systems | Coding-agent evaluation | AI safety and assurance | Responsible AI | Multilingual NLP
I build and evaluate agentic systems, with a focus on coding-agent controls and trace-level evaluation. My wider research covers responsible AI, multilingual NLP, low-resource African languages, and culturally grounded emotion modelling.
I am currently a PhD Researcher in Data Science and AI at the University of Hull, where my thesis explores cross-cultural musical elements, emotional expression, and genre characteristics in contemporary global music lyrics using NLP and deep learning. I was awarded a fully funded Faculty of Science and Engineering PhD Scholarship in Data Science. I also teach machine learning, NLP, and deep learning laboratory sessions to MSc students, supporting practical model development, Python engineering, and applied AI evaluation.
Website | Research portfolio | LinkedIn
- Agentic systems and security: Request authentication, tool permissions, approval controls, session-aware policies, and redacted audit records.
- Coding-agent evaluation and assurance: Trace-level benchmarks, evidence localisation, intervention timing, causal policy probes, and safety-utility measurement.
- Responsible AI and cultural data governance: Community-led governance, digital and AI literacy, bias-aware NLP, ethical dataset use, and locally responsible research practice.
- Low-resource and multilingual NLP: African language processing, code-mixed text, tokenisation, language identification, cross-lingual transfer, and evaluation for underrepresented languages.
- Emotion recognition and music AI: Fine-grained emotion modelling for multilingual lyrics, cultural context, class imbalance, and responsible interpretation of affective labels.
- Retrieval and research systems: Retrieval-augmented classification, evidence screening, provenance, structured outputs, and reproducible research workflows.
-
Secure Agent Gateway A Python gateway for authenticated agent tool requests, role and parameter policies, request-bound approvals, and redacted audit records. Version 0.5.0 adds sequence-aware controls, bounded policy checking, policy-change analysis, and causal-temporal probe synthesis.
-
Agentic Security Control Bench A DOI-archived contrastive benchmark for evaluating coding-agent monitors. Its v0.4.1 release contains a 320-trace benchmark and a separate 160-trace adversarial holdout, with measures for prevention, permitted-task retention, evidence grounding, intervention timing, calibration, and approval burden.
-
Low-Resource NLP Toolkit A released Python package for African language pre-processing, selective language routing, code-switch audits, emotion-label mapping, and coverage-aware evaluation. Its pinned AfriSenti benchmark covers 18,402 held-out test examples across five languages. Available on PyPI.
-
Multilingual DimStance Baselines Reproducible baselines for SemEval-2026 Task 3, Track B, with fixed data hashes, grouped out-of-fold evaluation, low-resource training budgets, and uncertainty intervals. Aspect conditioning reduced test macro RMSE from 1.343 to 1.323.
-
Coding Agent Failure Atlas A labelled synthetic trace dataset for coding-agent monitor research, with evidence spans, intervention points, and safer counterfactuals.
-
Coding Agent Monitor Lab An evaluation harness for testing whether monitors catch risky coding-agent traces using structured evidence and localised failure labels.
- Co-author and presenter, Digital Humanities 2025, Lisbon: locally responsible artificial intelligence frameworks for community-led digital data governance of cultural heritage in Burkina Faso.
- Co-author, Detection of Persuasion in Memes Across Languages with Ensemble Learning and External Knowledge: multilingual persuasion detection work in the ACL/SemEval shared-task research space.
- Research contributor on British Academy ODA and UNESCO capacity-building work for responsible technology use, AI literacy, intellectual property, digital visibility, and intangible cultural heritage practitioners.
- PhD research on multilingual emotion recognition, retrieval-augmented classification, African language processing, and culturally grounded music emotion modelling.
- UKRI Member, EPSRC and NERC Peer Review Colleges
- Reviewer, International Conference on Learning Representations
- Reviewer, The Deep Learning Indaba
- Associate Fellow, Advance HE
- Professional Member, BCS - The Chartered Institute for IT
- Panel Speaker, 6th European Chatbot and Conversational AI Summit, Edinburgh
- Co-author and Presenter, Digital Humanities 2025, Lisbon
- Invited Speaker, PyCon Lithuania 2025
- Lead Organiser and Speaker, Towards Transparent and Responsible AI Conference, University of Hull
- Invited Speaker, Pint of Science UK, Hull
- Featured Contributor, BBC News Interview for National AI Day
Languages and frameworks: Python, PyTorch, TensorFlow, Hugging Face Transformers, scikit-learn, LangChain, LangGraph, Bash
Agent systems and assurance: coding-agent evaluation, policy enforcement, approval workflows, audit logging, agent security, and model evaluation
AI and NLP: multilingual NLP, low-resource language processing, emotion classification, model fine-tuning, RAG, and bias and fairness metrics
Data and engineering: pandas, NumPy, reproducible research tooling, structured JSON/CSV outputs, Git, Docker, Azure
Human languages: English, Yoruba, Spanish, French
- Website: oyinkanchekwas.com
- Research portfolio: oyinkanchekwas.github.io
- LinkedIn: linkedin.com/in/oyinkan-chekwas
- GitHub: github.com/oyinkanchekwas


