Shaluka Gimhan

Shaluka Gimhan sh4lu-z

I like learning about deep neural nets.


2025 -

G.C.E. Advanced Level & Open Source

Studying A/L Physical Science while building open-source projects and tech communities.

  • G.C.E. Advanced Level: Studying A/L Physical Science.
  • Syntiox Lab: Exploring AI fundamentals and experimenting with LLMs.
  • Syntiox: Leading an open-source initiative for developers.
2019 - 2025

G.C.E. Ordinary Level (O/L)

Successfully completes the GCE Ordinary Level examinations.

  • Web Development: Participated in the Cyber Lowata Piyapath project, developing functional websites for schools.
  • Programming Concepts: Gained basic programming and markup knowledge using HTML and Pascal.

About

I am Shaluka Gimhan, a student and I am interested in neural nets, and I love building and working with AI & cyber security.I'm a fast learner and I can work well with others. I like to build new things from what I learn.


Expertise

Cyber Security

Advanced penetration testing and ethical hacking tools designed for modern digital threats.

AI & NLP

Pre-training and fine-tuning LLMs with custom tokenizers for high-performance language modeling.

Open Source

Robust systems and networking tools built using Rust and Node.js for maximum efficiency.


Models & Datasets

ceylex_x

A custom-engineered artificial intelligence model by Syntiox. Specializes in advanced text comprehension and structural execution natively within its weights.

Hugging Face • ◿ 5B • ⤓ 5+

[Closed Source Model]

This is a proprietary, closed-source language model optimized for internal business logic and data security.

Private License • Confidential

Sinhala-Qwen3-v7500

A Qwen3-based language model fine-tuned for Sinhala text generation with a vocabulary of 7,500 tokens.

Hugging Face • ◿ 2B • ⤓ 1.54k

Sinhala-Mega-Corpus-v1

A large-scale Sinhala text dataset built for training and evaluating NLP models on the Sinhala language.

Hugging Face • ☰ 323k • ⤓ 21

awesome-dataset-sinhala

A curated Sinhala language dataset covering diverse text domains for language modeling.

Hugging Face • ☰ 1.08M • ⤓ 9

CC100-sinhala

A Sinhala subset of the CC-100 multilingual dataset extracted from CommonCrawl for pretraining.

Hugging Face • ☰ 12.6M • ⤓ 13

Directory

Projects

A collection of my open-source tools and technical projects.

Tools ↗

multi-functional toolkit of 90+ web-based utilities, including PDF management, image processing, and educational resources for everyday use.


Connect

Find my work across different platforms or get in touch for collaborations.

View Profiles & Contact →