🚀 معرض المشاريع · 2026Project Portfolio · 2026

ضياء الدين السبيعيDhiyaa Aldeen Al-Subaie

DHIYAA ALDEEN AL‑SUBAIEGEMNEX AI · FOUNDER & AI ENGINEER
🔒ذكاء اصطناعي على جهازك · لا شيء يغادرهOn-device AI · nothing leaves the machine
🧩رؤية · كلام · نماذج لغوية · وكلاء واجهاتVision · Speech · LLMs · GUI agents
🏗️٨ أنظمة · فردياً · من الصفر8 systems · solo · from scratch
🗂️ أنظمة مختارةSelected systems

المشاريع — مرتّبة حسب التأثيرProjects — ordered by impact

👁️01NUWANUWA
رؤية حاسوبية · على الجهازComputer Vision · On-Premتوسيم وتدريب رؤية حاسوبية على الجهاز — دورة كاملة، ولا شيء يغادر الجهازOn-prem CV labeling & training — full lifecycle, nothing leaves the machine
Grounding DINOVLMYOLOPyTorch
🎙️02زينةZeina
ذكاء محادثة · ثلاثي الأبعادConversational AI · 3Dمحاوِرة ذكاء اصطناعي ثلاثية الأبعاد تُجري مقابلة أولى حقيقية لكل مرشّح3D AI interviewer that runs a real first round for every candidate
Whisper ×2Groq LPUWebSocketThree.js
🧍03التوأمTwin
رؤية حاسوبية · لحظيComputer Vision · Realtimeتوأم رقمي حيّ ثلاثي الأبعاد يتعرّف على الأشخاص ويتتبّعهم عبر الكاميراتLive 3D digital twin that recognizes and tracks people across cameras
InsightFaceYOLO26FAISSEmbeddings
04SIGHTSIGHT
وكيل واجهات · إدراكGUI Agent · Perceptionعين وكيل واجهات فائقة السرعة — إدراك في ~0.4 ثانية، مبنيّ من الصفرUltra-fast GUI-agent vision — perception in ~0.4s, built from scratch
YOLO26CRNN+CTCCNNOpenCV
📊05AI / MLAI / ML
تحليلات · تعلّم آليAnalytics · AutoMLتحليلات وتعلّم آلي مكتبي — من ملف خام إلى نموذج موثوق، بلا كود ولا سحابةDesktop analytics & AutoML — raw file to trusted model, no code, no cloud
XGBoostCUDAEnsembleFlask
🎧06الكلامSpeech
كلام · مصنع بياناتSpeech · Data Engineخط آلي لإنتاج بيانات تدريب عربية عالية الجودة لتحويل الكلام إلى نصAutomated pipeline for high-quality Arabic speech-to-text training data
3× ASRWhisperXDiarizationCloud
🔬07LLMLLM
نماذج لغوية · بحث وتطويرLLM · R&Dنموذج لغوي مبنيّ من الصفر — إثبات عمق فهم البنية كاملةًAn LLM built from scratch — proof of end-to-end architecture depth
TransformerCustom BPEB200/H100Cerebras
🔎08RAGRAG
استرجاع · بحث وتطويرRetrieval · R&Dاسترجاع متقدّم بثلاث طبقات مع تجارب حقيقية في التعلّم المعزّز والتضمينAdvanced three-layer retrieval with real RL and embedding experiments
Hybrid RAGASTQwen3BM25
SYS-01رؤية حاسوبية · على الجهازComputer Vision · On-Prem
01

👁️ حين يكون السؤال «أيهما مسموح أصلاً» لا «أيهما أرخص»👁️ When the question is "which is even allowed" — not "which is cheaper"

منصة توسيم وتدريب رؤية حاسوبية تعمل بالكامل داخل شبكة العميل، تغطي دورة الحياة الكاملة: جمع البيانات ← التوسيم الآلي ← التحقق ← التدريب ← التقييم ← إعادة التوسيم بالنموذج المُدرَّب. لا تُخرج صورة أو توسيماً خارج الجهاز — الخصوصية خاصية معمارية قابلة للتدقيق، لا بند في سياسة.A computer-vision labeling & training platform that runs entirely inside the client's network, covering the full lifecycle: data collection → auto-labeling → verification → training → evaluation → re-labeling with the trained model. No image or annotation ever leaves the machine — privacy is an auditable architectural property, not a policy clause.

🧰 التقنياتStackGrounding DINOVLMYOLOPyTorchCUDAMultiprocess
🖼️
نظام NUWA — الواجهة الرئيسيةNUWA Labeling — main interface
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · نظام NUWA — الواجهة الرئيسيةFIG · NUWA Labeling — main interface
🎬 شاهد ديمو NUWA Labeling🎬 Watch the NUWA Labeling demo
🔊 يعمل هنا مباشرة — مع الصوتPlays right here — with sound · فتح على YouTubeopen on YouTube
⚠️ المشكلةThe problem

شركات الرؤية الحاسوبية تعمل على لقطات CCTV وبيانات حسّاسة، ورفعها لمنصة سحابية عائق تعاقدي وتنظيمي — لا مجرد تفضيل. ثم يأتي ألم البداية من الصفر (cold start): توسيم يدوي يكلّف أياماً قبل أن يرى النموذج أول صورة، وتكلفة لكل مقعد ولكل صورة تجعل المشروع الكبير غير مجدٍ أصلاً.Computer-vision teams work on CCTV footage and sensitive data, and uploading it to a cloud platform is a contractual and regulatory blocker — not just a preference. Then comes the cold-start pain: manual labeling costs days before the model sees its first image, and per-seat, per-image pricing makes large projects unviable to begin with.

الحلThe solution

منصة on-premises تغطي الدورة كاملة محلياً: التوسيم الآلي يبدأ بلا صورة موسومة واحدة عبر Grounding DINO، فيتقلّص دور الإنسان من الرسم إلى الموافقة. التكلفة الحدّية بعد التركيب تقارب الصفر، وعلى المجموعات الكبيرة يتحوّل هذا من توفير هامشي إلى فارق في جدوى المشروع نفسه.An on-premises platform covering the full cycle locally: auto-labeling starts with zero labeled images via Grounding DINO, shrinking the human's role from drawing to approving. Marginal cost after install is near zero — and on large datasets this shifts from marginal savings to the difference in whether the project is viable at all.

🛠️ التحدّي الهندسي الذي حسم النضجThe engineering challenge that decided maturity

المشكلة الجوهرية في أدوات التوسيم المكتبية: استدلال PyTorch يمسك الـ GIL فيتجمّد التطبيق أثناء العمل الثقيل. الحل المطبَّق هنا يفصل ثلاث طبقات تزامن — واجهة، خيوط إشراف، وعمليات استدلال معزولة بذاكرتها وسياق CUDA خاص بها — فتبقى الواجهة حيّة تحت أثقل حمل. هذه الصلابة الهندسية هي ما حوّلها من نموذج أوّلي إلى منصة إنتاجية (~25 ألف سطر) بُني عليها 10 منتجات تُستخدم تجارياً وتدرّ دخلاً.The core problem in desktop labeling tools: PyTorch inference holds the GIL, freezing the app during heavy work. The solution here separates three concurrency layers — UI, supervisor threads, and isolated inference processes each with their own memory and CUDA context — so the interface stays alive under the heaviest load. That engineering robustness is what turned it from a prototype into a production platform (~25K lines) on top of which 10 commercially-used, revenue-generating products were built.

🧩 تحت الغطاءUnder the hood
توسيم آلي من مرحلتينTwo-stage auto-labeling
Grounding DINO للكشف zero-shot بلا صورة موسومة واحدة، ثم تحقّق آلي بنموذج لغوي بصري (VLM) — دور الإنسان يتقلّص من الرسم إلى الموافقة.Grounding DINO for zero-shot detection with no labeled image, then automatic verification by a vision-language model (VLM) — the human's role shrinks from drawing to approving.
حلقة مغلقة تتحسّنSelf-improving closed loop
النموذج المُدرَّب على الدفعة الأولى يوسّم الدفعة التالية، فتتحسّن الاقتراحات وتنخفض كلفة التوسيم مع كل دورة.The model trained on the first batch labels the next one, so suggestions improve and labeling cost drops with every cycle.
تدريب YOLO من نفس الشاشةYOLO training in-app
كشف · تقسيم · تصنيف، بلوحة مقاييس حية (mAP, Precision, Recall)، مراقبة VRAM، واستئناف من نقطة تفتيش.Detection · segmentation · classification, with a live metrics panel (mAP, Precision, Recall), VRAM monitoring, and checkpoint resume.
تقييم ومقارنة نسخEvaluation & model comparison
مصفوفات ارتباك ومنحنيات PR/F1 حية، ومقارنة نماذج جنباً إلى جنب — تحوّل الأداة من «موسِّم» إلى منصة تجريب.Live confusion matrices and PR/F1 curves, plus side-by-side model comparison — turning the tool from a "labeler" into an experimentation platform.
10
منتجات مبنية عليهاproducts built on it
25K+
سطر برمجي · 42 صنفاًlines of code · 42 classes
3
طبقات تزامن معزولةisolated concurrency layers
~0
تكلفة حدّية بعد التركيبmarginal cost after install
🔒 الشريحة المستهدفة الأوضح: فرق الرؤية الحاسوبية على بيانات CCTV وبيانات خاضعة لقيود — حيث السؤال ليس «أيهما أرخص» بل «أيهما مسموح أصلاً».🔒 The clearest target segment: computer-vision teams on CCTV and restricted data — where the question isn't "which is cheaper" but "which is even allowed".
SYS-02ذكاء محادثة · ثلاثي الأبعادConversational AI · 3D
02

🎙️ «زينة» — كل مرشّح يأخذ مقابلة حقيقية، لا رقم في طابور🎙️ "Zeina" — every candidate gets a real interview, not a number in a queue

محاوِرة افتراضية ثلاثية الأبعاد تتكلم بلهجة سعودية، تدير المقابلة الأولى كاملة لكل مرشّح بجودة ثابتة، وتُسلّم تقرير تقييم فور انتهائها. النتيجة: فريق التوظيف يبدأ من قائمة مختصرة مسنودة بحوار فعلي — لا بقراءة ست ثوانٍ لسيرة ذاتية.A 3D virtual interviewer that speaks in a Saudi dialect, runs the full first-round interview for every candidate at consistent quality, and delivers an assessment report the moment it ends. The result: the hiring team starts from a shortlist backed by an actual conversation — not a six-second résumé skim.

🧰 التقنياتStackWhisper ×2Groq LPUWebSocketThree.js3D Avatar
🖼️
زينة — واجهة المُحاوِرZeina — interviewer interface
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · زينة — واجهة المُحاوِرFIG · Zeina — interviewer interface
🎬 شاهد ديمو «زينة» الكامل🎬 Watch the full "Zeina" demo
🔊 يعمل هنا مباشرة — مع الصوتPlays right here — with sound · فتح على YouTubeopen on YouTube
⚠️ المشكلةThe problem

مئات السِيَر تدخل، وطاقة الفريق البشري لا تتجاوز عدداً محدوداً من المقابلات. النتيجة ثلاث خسائر في آنٍ واحد: مرشّحون أكفاء يُستبعدون لأن أحداً لم يفتح ملفهم، وأسابيع تُهدر في مقابلات أولى متكررة على نفس الأسئلة، ومرشّح ينتظر بلا رد فتخسر الشركة سمعتها كجهة توظيف.Hundreds of résumés come in, and a human team can only run a limited number of interviews. The result is three losses at once: strong candidates dropped because no one opened their file, weeks wasted on repetitive first-round interviews asking the same questions, and candidates left waiting — costing the company its employer reputation.

الحلThe solution

مقابلة أولى آلية لكل مرشّح بلا استثناء وبنفس الجودة: حوار طبيعي، أسئلة مبنية على سيرته والوظيفة، وتقرير تقييم فوري. يتحوّل عمل الفريق البشري من غربلة مئات الملفات إلى اتخاذ قرار على قائمة قصيرة موثوقة — أسابيع تنكمش إلى ساعات.An automated first-round interview for every candidate, without exception, at the same quality: a natural conversation, questions grounded in their CV and the role, and an instant assessment report. The human team's job shifts from screening hundreds of files to deciding over a trusted shortlist — weeks shrink to hours.

🛠️ لماذا هذه الهندسة تحديداًWhy this exact architecture

مقابلة حقيقية لا تحتمل تأخيراً محرجاً ولا هلوسة تُهين المرشّح. فبنيت كل طبقة لخدمة إحساس واحد: أنك تحاور إنساناً. لذلك نموذجا Whisper يعملان بالتوازي حتى لا يُمضغ لهجتك أو مصطلحك، والبث عبر WebSocket بـ ~150ms حتى لا تنتظر ردّها، والاستدلال على Groq LPU حتى يفكّر السؤال التالي أسرع مما تُكمل جوابك. والأهم — طبقة عدالة تضمن أن يُقيَّم كل مرشّح على ما قاله فعلاً، لا على خطأ في تحويل الصوت. جرّبها متطوّعون حقيقيون، وجرّبتها بنفسي: التجربة ممتعة والمستوى فخم فعلاً.A real interview can't tolerate an awkward delay or a hallucination that insults the candidate. So every layer was built to serve one feeling: that you're talking to a human. Two Whisper models run in parallel so your dialect or technical term isn't garbled; streaming over WebSocket at ~150ms so you don't wait for its reply; inference on Groq LPU so the next question is thought up faster than you finish answering. Most important — a fairness layer ensures each candidate is judged on what they actually said, not on a speech-to-text error. Real volunteers tried it, and so did I: the experience is genuinely enjoyable and the quality feels premium.

🧭 كيف تبدو التجربة للمرشّح؟What the candidate experience looks like
1

تفتح الرابط — بلا تحميل تطبيق ولا تسجيلOpen the link — no app install, no signup

2

زينة ترحّب بك بصوت طبيعي وتبدأ الحوارZeina greets you in a natural voice and starts the conversation

3

كلامك يتحول لنص وأنت تتكلم — لحظياًYour speech turns to text as you talk — in real time

4

أسئلة من سيرتك أنت والوظيفة، تتعمّق مع كل إجابةQuestions drawn from your CV and the role, deepening with each answer

5

أجندة حية تريك ما غُطّي وما تبقّى — تستدرك بنفسكA live agenda shows what's covered and what's left — it self-corrects

6

تقرير تقييمك كاملاً لحظة انتهاء المقابلةYour full assessment report the moment the interview ends

🧩 تحت الغطاءUnder the hood
نموذجا Whisper بالتوازيTwo Whisper models in parallel
عربي وإنجليزي مع طبقة دمج تقيس ثقة كل نموذج لكل مقطع وتختار الأدق — فتُلتقط اللهجة الخليجية والمصطلح التقني معاً بدل أن يمضغ نموذج واحد إحداهما.Arabic and English with a fusion layer that scores each model's confidence per segment and picks the most accurate — so the Gulf dialect and the technical term are both captured, instead of one model garbling either.
بث صوتي عبر WebSocketAudio streaming over WebSocket
زمن استجابة ~150ms — الصوت يتدفق مباشرة للمحرك بلا رفع ملفات ولا انتظار، وهو ما يجعل الحوار يبدو آنياً.~150ms latency — audio streams straight to the engine with no file uploads and no waiting, which is what makes the conversation feel instant.
استدلال على Groq LPUInference on Groq LPU
منصة معجّلة بالعتاد تولّد السؤال التالي شبه فورياً — بلا الصمتة المحرجة التي تكشف أنك تكلّم آلة.A hardware-accelerated platform generates the next question almost instantly — without the awkward silence that reveals you're talking to a machine.
أفاتار ثلاثي الأبعاد3D avatar
مزامنة شفاه وتعابير طبيعية عبر WebGL / Three.js — وجه يحاورك، لا شاشة نص باردة.Lip-sync and natural expressions via WebGL / Three.js — a face that talks with you, not a cold text screen.
محرك أسئلة تكيّفيAdaptive question engine
يبني أجندة من الوصف الوظيفي، يتتبع تغطيتها حياً، ويمنع الأسئلة القالبية — كل سؤال مرسّخ في شيء قاله المرشّح أو ورد في سيرته.Builds an agenda from the job description, tracks its coverage live, and blocks template questions — every question is anchored in something the candidate said or something in their CV.
بنية مقاومة للأعطالFault-tolerant architecture
كل محرك له بديل يسقط إليه تلقائياً، فالمقابلة لا تتوقف مهما حصل. وللشركات: رفع دفعة سِيَر ← رابط مقابلة لكل مرشّح تلقائياً + لوحة متابعة حية.Every engine has an automatic fallback, so the interview never stops whatever happens. For companies: upload a batch of CVs → an interview link per candidate automatically + a live tracking dashboard.
SYS-03رؤية حاسوبية · لحظيComputer Vision · Realtime
03

🧍 توأم رقمي حيّ — يحوّل كل شخص في المكان إلى شخصية ثلاثية الأبعاد تتحرّك معه🧍 A live digital twin — turning everyone in the space into a 3D character that moves with them

نظام رؤية حاسوبية يبني توأماً رقمياً للبيئة لحظياً: يكتشف كل شخص داخل منطقة محددة، يحوّله إلى شخصية 3D تتبع حركته الفعلية، يتعرّف عليه بوجهه عبر كاميرات CCTV متعددة، ويعرض فوق شخصيته اسمه ورتبته وصورته الملتقطة آنياً.A computer-vision system that builds a real-time digital twin of the environment: it detects every person in a defined area, turns them into a 3D character that follows their actual movement, recognizes them by face across multiple CCTV cameras, and displays their name, rank and a live-captured photo above their character.

🧰 التقنياتStackInsightFaceYOLO26FAISSEmbeddingsMulti-cam
🖼️
التوأم الرقمي — عرض حيّDigital twin — live view
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · التوأم الرقمي — عرض حيّFIG · Digital twin — live view
⚠️ المشكلةThe problem

كاميرات CCTV التقليدية تسجيل سلبي بلا فهم: لا تعرف مَن الشخص، تفقد أثره حين يستدير أو يمرّ خلف عائق، ولا تربط اللقطات بين كاميرات متعددة. النتيجة مراقبة بلا وعي لحظي بمن هو موجود وأين، وحضور يُرصد يدوياً.Traditional CCTV is passive recording with no understanding: it doesn't know who the person is, loses track when they turn or pass behind an obstacle, and never links footage across cameras. The result is surveillance with no live awareness of who is present and where, and attendance tracked by hand.

الحلThe solution

توأم رقمي 3D يعيد بناء المشهد حياً: كل شخص شخصية تتحرك بحركته، معرّفة بالاسم والرتبة وصورة آنية فوق رأسه، ومتتبَّعة عبر عدة كاميرات دون فقدان الأثر. أُثبت عملياً على أكثر من 10 أشخاص في أوقات مختلفة بأداء ممتاز — ويصلح كذلك كنظام حضور تلقائي فعّال.A 3D digital twin that rebuilds the scene live: each person is a character that moves as they move, tagged with name, rank and a live photo above their head, and tracked across multiple cameras without losing the trail. Proven in practice on 10+ people at different times with excellent performance — and it also works as an effective automatic attendance system.

🛠️ لماذا هذه الهندسة — معركة عدم فقدان الأثرWhy this architecture — the fight against losing the trail

أصعب تحدٍّ في التتبع هو فقدان الوجه: يختفي حين يستدير الشخص أو يبتعد فينكسر التعرّف. لذلك بنيت مساراً هجيناً متيناً — InsightFace يتعرّف بالوجه أولاً، فإذا فُقد الوجه يواصل التتبع عبر صندوق الجسم من YOLO nano (26) فلا ينقطع الأثر. وللتعرّف القوي بدل صورة واحدة: عدة لقطات لكل شخص + augmentation، تُخزَّن كـ embeddings وتُسترجع عبر FAISS بسرعة ودقة عالية حتى مع اختلاف الإضاءة والزاوية. هذه التركيبة تحديداً هي ما جعل الأداء ممتازاً عبر كاميرات وأوقات مختلفة.The hardest tracking challenge is losing the face: it disappears when the person turns or moves away, breaking recognition. So I built a robust hybrid path — InsightFace recognizes the face first, and if the face is lost it keeps tracking via the body box from YOLO nano (26) so the trail never breaks. For robust recognition, instead of a single image: several shots per person + augmentation, stored as embeddings and retrieved via FAISS quickly and accurately even under different lighting and angles. That exact combination is what made performance excellent across cameras and times.

🧩 تحت الغطاءUnder the hood
توأم رقمي 3D حيّLive 3D digital twin
كل شخص شخصية ثلاثية الأبعاد تتبع حركته الفعلية داخل المنطقة لحظياً — لا نقاط على خريطة، بل مشهد مُعاد بناؤه.Each person is a 3D character following their actual movement in the area in real time — not dots on a map, but a reconstructed scene.
تعرّف بالوجه متعدد الكاميراتMulti-camera face recognition
InsightFace يميّز الأشخاص عبر عدة كاميرات CCTV ويربط الهوية نفسها بينها — لا تكرار ولا التباس.InsightFace identifies people across several CCTV cameras and links the same identity between them — no duplication, no confusion.
تتبّع مقاوم للفقدانLoss-resistant tracking
عند اختفاء الوجه يكمل التتبع عبر صندوق الجسم (YOLO nano 26) — الأثر لا ينكسر حتى مع الاستدارة أو الحجب.When the face disappears, tracking continues via the body box (YOLO nano 26) — the trail doesn't break even with turning or occlusion.
بصمات قوية عبر FAISSStrong embeddings via FAISS
عدة لقطات لكل شخص + augmentation، تُخزَّن embeddings وتُسترجع بسرعة ودقة عالية — متانة أمام الإضاءة والزاوية.Several shots per person + augmentation, stored as embeddings and retrieved fast and accurately — robust to lighting and angle.
رتب وصور آنية فوق الشخصيةRanks & live photos above the character
يعرض الاسم والرتبة الحقيقية (موظف · مميّز · CEO) وصورة الشخص الملتقطة من الوقت الفعلي فوق توأمه الرقمي.Shows the real name and rank (Employee · VIP · CEO) and a live-captured photo of the person above their digital twin.
إيماءات اليد للتحكّم بالحالةHand gestures to control status
رفع اليد فوق الرأس يضبطك «مشغول» ويطلق إشعاراً في الواجهة؛ تكرار الإيماءة يزيلها، وترتفع تلقائياً بعد ~ساعة.Raising a hand above the head sets you to "busy" and fires a notification in the UI; repeating the gesture clears it, and it auto-clears after ~an hour.
🛠️ تجربة الإيماءات — ما نجح فعلاًThe gesture experiment — what actually worked

جرّبت عدة إيماءات: إشارة «OK» تعطي «تم»، لكنها أضعف على المسافات البعيدة. أما رفع اليد فوق الرأس فكان ممتازاً حتى من بعيد، فاعتمدته لضبط حالة «مشغول» مع إشعار يُسجَّل في الواجهة (تعمل على شاشة المكتب). ترتفع الحالة تلقائياً بعد نحو ساعة أو بتكرار الإيماءة. وأثبت النظام فاعليته أيضاً في بيئة الحضور التلقائي.I tried several gestures: an "OK" sign for "done", but it was weaker at long range. Raising a hand above the head worked excellently even from afar, so I adopted it to set "busy" status with a notification logged in the UI (runs on a desktop screen). The status auto-clears after about an hour or by repeating the gesture. The system also proved effective in an automatic attendance setting.

10+
أشخاص مختبَرون بنجاحpeople tested successfully
CCTV
كاميرات متعددة مترابطةmultiple linked cameras
3D
توأم حيّ يتبع الحركةlive twin following movement
حضور تلقائي فعّالeffective auto-attendance
SYS-04وكيل واجهات · إدراكGUI Agent · Perception
04

⚡ عين وكيل يرى الشاشة ويتحكّم بها — بإيقاع تفاعلي لا انتظار⚡ An agent's eye that sees and controls the screen — at an interactive pace, no waiting

المنتج ليس OCR — بل وكيل يرى الشاشة، يفهمها، ويتحكّم بها. هذا المحرّك هو العين: طبقة إدراك (OCR + كشف عناصر الواجهة) مدرَّبة بالكامل من الصفر، تعمل محلياً في ~0.4 ثانية، وتغذّي دماغاً لغوياً على Cerebras — فتصير الحلقة الكاملة في نطاق أقل من ثانية للخطوة.The product isn't OCR — it's an agent that sees the screen, understands it, and controls it. This engine is the eye: a perception layer (OCR + UI-element detection) trained entirely from scratch, running locally in ~0.4s, feeding a language brain on Cerebras — so the full loop lands in the range of under one second per step.

🧰 التقنياتStackYOLO26CRNN+CTCCNNOpenCVCerebras
🖼️
GEMNEX SIGHT — واجهة الوكيلGEMNEX SIGHT — agent view
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · GEMNEX SIGHT — واجهة الوكيلFIG · GEMNEX SIGHT — agent view
⚠️ المشكلةThe problem

وكلاء الواجهات اليوم بطيئون، والبطء من طرفين: الإدراك (محرّكات OCR عامة أو نماذج VLM ضخمة تأخذ ثوانٍ لكل لقطة) والتخطيط (نماذج لغوية على بنية تقليدية). والحلقة تسير بسرعة أبطأ مكوّن فيها — فتسريع طرف واحد لا يعطي شيئاً، والنتيجة وكيل «انتظر ثم شاهد» لا وكيل يعمل.GUI agents today are slow, and the slowness comes from two ends: perception (general OCR engines or huge VLMs taking seconds per frame) and planning (language models on conventional hardware). The loop runs at the speed of its slowest component — so speeding up one end gives you nothing, and you get a "wait-then-watch" agent, not an agent that works.

الحلThe solution

عالجنا الطرفين معاً: إدراك محلّي مخصّص في ~0.4 ثانية + تخطيط فائق السرعة على Cerebras. النتيجة انتقال من إيقاع «انتظر ثم شاهد» إلى تفاعل آنيّ حقيقي — ووضع خصوصية صارم: العين لا ترسل بكسلاً واحداً خارج الجهاز، والنموذج اللغوي يستقبل نصاً وإحداثيات فقط لا صورة الشاشة.We addressed both ends together: custom local perception in ~0.4s + ultra-fast planning on Cerebras. The result is a shift from "wait-then-watch" to true real-time interaction — with a strict privacy mode: the eye sends not a single pixel off-device, and the language model receives only text and coordinates, never the screen image.

🛠️ لماذا بنيت نظام OCR من الصفرWhy I built an OCR system from scratch

جرّبت المحرّكات العالمية أولاً — EasyOCR وPaddleOCR وRapidOCR — فوجدتها تأخذ ثوانيَ على اللقطة الواحدة، وهو ما يجعل بناء وكيل تفاعلي فوقها مستحيلاً عملياً. لقطة الواجهة نوع مختلف كليّاً عن صور المستندات: نصّ صغير جداً، كثافة مئات الصناديق، أيقونات ذات معنى، وعربية وإنجليزية تختلطان في السطر. فقرّرت التدريب من الصفر: كاشف YOLO26 مضبوط على كثافة الواجهات، متعرّف CRNN+CTC بمعالجة عربية دقيقة، وخطّ استدلال بأشكال ثابتة — حتى نزل الزمن من 5 ثوانٍ إلى ~0.4. كل ذلك لسبب واحد: الوصول إلى وكيل سريع جداً وبأداء قوي.I tried the mainstream engines first — EasyOCR, PaddleOCR, RapidOCR — and found they take seconds per frame, which makes building an interactive agent on top of them practically impossible. A UI screenshot is a completely different kind of image from document scans: tiny text, hundreds of dense boxes, meaningful icons, and Arabic and English mixed in one line. So I decided to train from scratch: a YOLO26 detector tuned for UI density, a CRNN+CTC recognizer with careful Arabic processing, and a fixed-shape inference pipeline — until latency dropped from 5 seconds to ~0.4. All for one reason: reaching a very fast, strong agent.

السرعة — مقارنة مُقاسة على نفس الصورة ونفس الجهازSpeed — measured on the same image, same machine
▸ الشريط الأقصر = أسرع▸ shorter bar = faster
⚡ GEMNEX SIGHT
≈0.4s · 1×
RapidOCR
16× slower
EasyOCR
32× slower

📐 كيف تُقرأ: قياس بعد تسخين كل محرّك، والرقم وسيط عدّة تشغيلات. جزء من الفارق أن نظامنا على GPU؛ لكن الفارق البنيوي الحقيقي من التخصّص: كاشف مضبوط على كثافة الواجهات، متعرّف بعرض قصير ملائم لأسطرها، وخطّ استدلال بأشكال ثابتة مسخّنة مسبقاً. عند ~0.4 ثانية للخطوة يصير التفاعل الآني ممكناً — وعند عدّة ثوانٍ لا يصير. هذا فارق في رتبة الحجم.📐 How to read it: measured after warming up each engine, with the median of several runs. Part of the gap is that our system is on GPU; but the real structural gap is from specialization: a detector tuned for UI density, a recognizer with a short width suited to UI lines, and a fixed-shape, pre-warmed inference pipeline. At ~0.4s per step, real-time interaction becomes possible — at several seconds, it doesn't. This is an order-of-magnitude difference.

🧩 تحت الغطاءUnder the hood
الكاشف — YOLO26The detector — YOLO26
يلتقط مئات الصناديق النصّية عالية الكثافة، ويميّز الأيقونات والإيموجي والعناصر البنيوية كأصناف مستقلة. جاهز عملياً.Captures hundreds of high-density text boxes and distinguishes icons, emojis and structural elements as separate classes. Production-ready.
المتعرّف — CRNN + CTCThe recognizer — CRNN + CTC
يقرأ العربية والإنجليزية بمعالجة خاصة: NFKC، الترتيب البصري RTL، وحذف التطويل — سبب فشل أغلب تدريبات CRNN العربية.Reads Arabic and English with special processing: NFKC, RTL visual ordering, and tatweel removal — the reason most Arabic CRNN trainings fail.
مصنّف أيقونات CNNCNN icon classifier
شبكة صغيرة ~0.35M بارامتر تعمل بسرعة حتى على المعالج — إضافة صنف جديد تكلّف ~20 عيّنة فقط.A small ~0.35M-parameter network runs fast even on CPU — adding a new class costs only ~20 samples.
من 5 ثوانٍ إلى < 0.5From 5 seconds to < 0.5
دلاء عرض ثابتة مسخّنة مسبقاً، معالجة OpenCV بدل PIL، وتصنيف الأيقونات داخل الكاشف نفسه — هندسة استدلال دقيقة أسقطت الزمن رتبة كاملة.Pre-warmed fixed width buckets, OpenCV processing instead of PIL, and icon classification inside the detector itself — precise inference engineering that dropped latency by a full order of magnitude.
🧠 الحلقة الكاملة: لقطة → إدراك (~0.4s) → تخطيط على Cerebras → تنفيذ نقر وكتابة — انتقال من إيقاع «انتظر ثم شاهد» إلى تفاعل آني حقيقي.🧠 The full loop: capture → perception (~0.4s) → planning on Cerebras → click & type execution — a shift from "wait-then-watch" to true real-time interaction.
SYS-05تحليلات · تعلّم آليAnalytics · AutoML
05

📊 ما كان يأخذ فريق بيانات أسابيع — يصير دقائق على جهازك📊 What used to take a data team weeks — now minutes on your own machine

تطبيق مكتبي يعمل بالكامل على جهاز العميل: يحوّل ملف بيانات خام (CSV / Excel / JSON) إلى لوحة تحليل بصرية + نموذج تعلّم آلي مدرَّب + توقّعات مستقبلية + ملف نموذج جاهز للنشر — دون سطر برمجي واحد، ودون أن يغادر بايت واحد جهازك.A desktop app that runs entirely on the client's machine: it turns a raw data file (CSV / Excel / JSON) into a visual analytics dashboard + a trained ML model + future forecasts + a deployment-ready model file — without a single line of code, and without a single byte leaving your machine.

🧰 التقنياتStackXGBoostCUDAEnsembleFlaskDesktop
🖼️
GEMNEX AI/ML — لوحة التحليلاتGEMNEX AI/ML — dashboard
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · GEMNEX AI/ML — لوحة التحليلاتFIG · GEMNEX AI/ML — dashboard
🎬 شاهد ديمو GEMNEX AI/ML🎬 Watch the GEMNEX AI/ML demo
🔊 يعمل هنا مباشرة — مع الصوتPlays right here — with sound · فتح على YouTubeopen on YouTube
⚠️ المشكلةThe problem

كل تحليل يتطلّب مختص Python/BI متفرّغاً، وكل سؤال بسيط يتحوّل إلى دورة طلب ← انتظار ← تسليم تمتد أياماً. أدوات السحابة مرفوضة أصلاً في القطاعات الحسّاسة، والأخطر: تُعرض أرقام دقة مرتفعة بلا تدقيق، فتُبنى قرارات على نماذج مضلِّلة تنهار في الإنتاج.Every analysis needs a dedicated Python/BI specialist, and every simple question turns into a request → wait → deliver cycle lasting days. Cloud tools are outright rejected in sensitive sectors, and worse: high accuracy numbers are shown without scrutiny, so decisions get built on misleading models that collapse in production.

الحلThe solution

«ارفع ملفك، وخلال دقائق تحصل على المسار الكامل»: تنظيف تلقائي، لوحة رسوم فورية، عدّة نماذج مدرَّبة ومقارنة، توقّعات مستقبلية، وملف نموذج جاهز للنشر — بلا كود، وبلا سحابة. والاعتماد على مختص متفرّغ يتحوّل من إلزامي إلى اختياري."Upload your file, and within minutes you get the full pipeline": automatic cleaning, an instant charts dashboard, several trained and compared models, future forecasts, and a deployment-ready model file — no code, no cloud. Relying on a dedicated specialist shifts from mandatory to optional.

🛠️ كيف بُني — والفرق الحقيقيHow it was built — and the real difference

الخوارزميات ليست الفجوة — فهي متاحة ومجانية. الفجوة في المسار الكامل من الملف الخام إلى قرار موثوق، وفي الثقة بأن الرقم المعروض صادق. بُني على ~3 أشهر و14 دورة تحسين موثّقة، كل واحدة انطلقت من استخدام حقيقي كشف مشكلة حقيقية، وكل إصلاح وثّق السبب الجذري لا العَرَض. أي أداة تُخرج رقم دقة — القليل منها يخبرك متى يكون هذا الرقم كذبة، وهذه بالضبط ميزتها التنافسية.The algorithms aren't the gap — they're free and available. The gap is the full pipeline from raw file to a trustworthy decision, and the confidence that the number shown is honest. Built over ~3 months and 14 documented improvement cycles, each starting from real usage that exposed a real problem, and each fix documenting the root cause, not the symptom. Any tool outputs an accuracy number — few tell you when that number is a lie, and that is exactly its competitive edge.

🧩 تحت الغطاءUnder the hood
لوحة تحليل تفاعليةInteractive analytics dashboard
Power BI محلية: 4 رسوم متنوعة تلقائياً، تعديل في مكانه، 9 أنواع + مصفوفة ارتباط، وذكاء كثافة يتحوّل لخريطة حرارية فوق 4,000 نقطة.Local Power BI: 4 varied charts automatically, in-place editing, 9 types + a correlation matrix, and density intelligence that switches to a heatmap above 4,000 points.
تدريب النماذجModel training
6 خوارزميات + نماذج تصويت جماعية، تحسين تلقائي، وتسريع XGBoost على CUDA مع سقوط تلقائي للمعالج.6 algorithms + ensemble voting models, automatic tuning, and XGBoost acceleration on CUDA with automatic CPU fallback.
تنبؤ بثلاثة مستوياتThree-level forecasting
زر واحد «توقّع الآن»، أو آفاق مستقبلية (7 · 30 · 90) بنطاق عدم يقين وتواريخ حقيقية، أو وضع متقدم بالتحكم الكامل.One "predict now" button, or future horizons (7 · 30 · 90) with an uncertainty range and real dates, or an advanced mode with full control.
طبقة الثقة — الميزة الحقيقيةThe trust layer — the real edge
حارس تسرّب بيانات يكشف «دقة 100%» الوهمية، واستبدال تلقائي عند R² سالب (مقيس: من −17.15 إلى 0.67). النظام يقول الحقيقة حين تكون البيانات غير كافية.A data-leakage guard that detects illusory "100% accuracy", and automatic replacement on negative R² (measured: from −17.15 to 0.67). The system tells the truth when the data is insufficient.
+90%
توفير في المراحل التقنيةsaved on technical stages
200–700
ساعة عمل موفَّرة شهرياًwork hours saved monthly
50
مسار خادم · 25 نوع رسمserver routes · 25 chart types
0
اتصال خارجي — البيانات لا تغادر الجهازexternal connections — data never leaves the machine
SYS-06كلام · مصنع بياناتSpeech · Data Engine
06

🎧 مصنع بيانات تحويل الكلام إلى نص — حيث السوق العربي فارغ🎧 A speech-to-text data factory — where the Arabic market is empty

خط إنتاج آلي يحوّل الصوت الخام إلى بيانات تدريب ASR عالية الجودة: ثلاثة محكّمين من أقوى نماذج التعرّف على الكلام، تحديد عدد المتحدثين، وتوقيت كل كلمة — ثم مراجعة وتعديل بمزوّدين، ولا يمرّ للتدريب إلا ما تجاوزت دقته 70%.An automated pipeline that turns raw audio into high-quality ASR training data: three judges from the strongest speech-recognition models, speaker counting, and per-word timing — then review and correction by two providers, and only data above 70% accuracy passes to training.

🧰 التقنياتStack3× ASRWhisperXDiarizationCloud
🖼️
مصنع بيانات الكلام — الواجهةSpeech Data Engine — interface
📤 انقر أو أفلت صورة📤 Click or drop an image
شكل · مصنع بيانات الكلام — الواجهةFIG · Speech Data Engine — interface
🔗🔗 عرض المنشور على LinkedIn🔗 View the post on LinkedIn
⚠️ المشكلةThe problem

سوق نماذج تحويل الكلام إلى نص عليه طلب عالٍ — خصوصاً بالعربية — لكن البيانات عالية الجودة نادرة ومحدودة. إنتاجها يدوياً يحتاج فرق توسيم كاملة، ووقتاً طويلاً، وتكلفة مرتفعة تجعل بناء نموذج عربي جيد عائقاً بذاته.The speech-to-text market has high demand — especially in Arabic — but high-quality data is scarce and limited. Producing it by hand needs entire labeling teams, long timelines, and high costs that make building a good Arabic model a blocker in itself.

الحلThe solution

خط إنتاج يحوّل الصوت الخام إلى بيانات تدريب موثوقة بأقل تدخّل بشري — يرفع الجودة، يقلّص الاعتماد على الموظفين، ويملأ فجوة السوق العربي بنسبة عالية. المخرَج بيانات جاهزة للتدريب، لا مجرد تفريغ صوتي.A pipeline that turns raw audio into reliable training data with minimal human intervention — raising quality, reducing reliance on staff, and largely filling the Arabic market gap. The output is training-ready data, not just a transcript.

🛠️ كيف صُمّم خط الإنتاجHow the pipeline was designed

لا أعتمد على نموذج واحد قد يخطئ: ثلاثة محكّمين من أقوى نماذج ASR يحكّمون على النص معاً، فتُلتقط الكلمة الصحيحة حتى حين يتعثّر أحدها. طبقة تحديد المتحدثين من مكتبة متخصصة على HuggingFace تفصل من قال ماذا، وWhisperX يستخرج توقيت كل كلمة على حدة لا الجملة. ثم مزوّدان قويان: أحدهما يراجع والآخر يعدّل، ولا يُعتمَد للتدريب إلا ما تجاوزت دقته 70% — مع خيار مراجعة يدوية للأجزاء الأعقد. كل ذلك يعمل على خدمة سحابية قابلة للتوسّع.I don't rely on a single model that might err: three judges from the strongest ASR models arbitrate the text together, so the correct word is captured even when one stumbles. A speaker-diarization layer from a specialized HuggingFace library separates who said what, and WhisperX extracts per-word — not per-sentence — timing. Then two strong providers: one reviews and the other corrects, and only data above 70% accuracy is accepted for training — with a manual-review option for the trickiest parts. All of it runs on a scalable cloud service.

🧩 المكوّناتComponents
ثلاثة محكّمين ASRThree ASR judges
ثلاثة من أقوى نماذج التعرّف على الكلام يحكّمون على النص معاً — دقة أعلى واستقرار أكبر من أي نموذج منفرد.Three of the strongest speech-recognition models arbitrate the text together — higher accuracy and more stability than any single model.
تحديد عدد المتحدثينSpeaker counting
مكتبة متخصصة من HuggingFace (speaker diarization) تفصل الأصوات وتنسب كل مقطع لقائله.A specialized HuggingFace library (speaker diarization) separates voices and attributes each segment to its speaker.
WhisperX — توقيت كل كلمةWhisperX — per-word timing
استخراج توقيت على مستوى الكلمة لا الجملة — دقّة محاذاة تفتح استخدامات لا تتيحها التفريغات العادية.Word-level, not sentence-level, timing — alignment precision that unlocks uses ordinary transcripts can't.
مزوّدان: مراجعة + تعديلTwo providers: review + correct
مزوّد قوي يراجع النصوص وآخر يعدّلها — طبقة تدقيق مزدوجة قبل اعتماد أي عيّنة.One strong provider reviews the text and another corrects it — a double verification layer before any sample is accepted.
عتبة دقة > 70% + مراجعة يدوية>70% accuracy threshold + manual review
لا يمرّ للتدريب إلا ما تجاوزت دقته 70%، مع خيار مراجعة بشرية للأجزاء المعقّدة جداً — الجودة قبل الكمّية.Only data above 70% accuracy passes to training, with a human-review option for the very complex parts — quality over quantity.
تشغيل سحابي قابل للتوسّعScalable cloud execution
الخط بأكمله يعمل على خدمة سحابية، فيعالج كميات صوت كبيرة دون اختناق على جهاز واحد.The entire pipeline runs on a cloud service, processing large volumes of audio without bottlenecking on a single machine.
SYS-07نماذج لغوية · بحث وتطويرLLM · R&D🏅 إثبات خبرةexpertise proof
07

🔬 بناء نموذج LLM من الصفر — لأثبت أنني أفهم كل طبقة🔬 Building an LLM from scratch — to prove I understand every layer

مشروع بحثي بنيته من الصفر عبر عشرات التجارب على عتاد سحابي عملاق: نموذج Transformers بانتباه كامل وترميز BPE مخصّص للعربية والإنجليزية وبايثون، وصل نتائج جيدة نسبةً لحجمه على نطاق محدّد (sub-domain). ليس استدعاء API — بل بناءٌ يثبت فهم البنية بالكامل.An R&D project I built from scratch through dozens of experiments on massive cloud hardware: a full-attention Transformer with a custom BPE tokenizer for Arabic, English and Python, reaching good results relative to its size on a specific sub-domain. Not an API call — a build that proves complete understanding of the architecture.

🧰 التقنياتStackTransformerCustom BPEB200/H100CerebrasMoE
⚠️ التحدّيThe challenge

بناء نموذج لغوي كفؤ من الصفر — لا استدعاء API جاهز — على ميزانية عتاد محدودة تفرض تجارب سريعة ومركّزة، مع توكنايزر يخدم العربية والإنجليزية والبرمجة معاً في آنٍ واحد.Building a competent language model from scratch — not a ready API call — on a limited hardware budget that forces fast, focused experiments, with a tokenizer serving Arabic, English and code all at once.

النتيجةThe result

نموذج Transformers بانتباه كامل وترميز BPE مخصّص، وصل نقطة جيدة نسبةً لحجمه على نطاق محدّد بعد سلسلة تجارب ممنهجة بمؤشرات أداء وطباعة عيّنات — وخبرة عملية عميقة بآليات الانتباه وتقنيات التسريع الحديثة.A full-attention Transformer with a custom BPE tokenizer, reaching a good point relative to its size on a specific domain after a systematic series of experiments with performance metrics and sample printouts — plus deep hands-on experience with attention mechanisms and modern acceleration techniques.

🧩 عمق التجربة الهندسيةDepth of the engineering experiment
Transformers + Full AttentionTransformers + Full Attention
بعد تجارب موسّعة يبقى الانتباه الكامل الأفضل لأنه يركّز على كل كلمة فعلياً — وقد لا يأتي ما يتفوّق عليه بسهولة.After extensive experiments, full attention remains best because it genuinely focuses on every word — and it may not be easily surpassed.
توكنايزر BPE مخصّصCustom BPE tokenizer
مصمّم ليلائم البرمجة (Python) والعربية والإنجليزية معاً — بيانات مركّزة على هذه النطاقات الثلاثة فقط.Designed to fit programming (Python), Arabic and English together — data focused on these three domains only.
أحجام معمارية متعددةMultiple architecture sizes
تجارب على 124M · 1B · 2B، مع مؤشرات أداء وطباعة عيّنات لرصد الجودة أثناء التدريب.Experiments on 124M · 1B · 2B, with performance metrics and sample printouts to monitor quality during training.
عتاد سحابي عملاقMassive cloud hardware
تجارب سريعة (لضبط التكلفة) على B200 · H100 عبر منصات مثل RunPod و Colab.Fast experiments (to control cost) on B200 · H100 via platforms like RunPod and Colab.
تجارب آليات انتباهAttention-mechanism experiments
اختبار بدائل مثل RWKV و Linear Attention والانتباه الهرمي و Sparse Attention — ومقارنتها بالانتباه الكامل.Testing alternatives like RWKV, Linear Attention, hierarchical attention and Sparse Attention — comparing them to full attention.
خلفية في تقنيات التسريعBackground in acceleration techniques
مزيج الخبراء (MoE)، فكّ الترميز التخميني (نموذج صغير يسوّد ونموذج كبير يتحقّق)، وذاكرة المفاتيح/القيم (KV Cache).Mixture of Experts (MoE), speculative decoding (a small model drafts, a large one verifies), and key/value memory (KV Cache).
تجربة حقيقية على Cerebras فائق السرعةA real experiment on ultra-fast Cerebras

جرّبت Cerebras فعلياً — عتاد بمقياس الرقاقة (wafer-scale) أكبر بكثير من كروت الشاشة الحالية، فيصير التوليد أسرع بمراحل. النتيجة المقيسة: نحو 3000 توكن/ثانية لنموذج بحجم 120B من OpenAI — سرعة يصعب حتى على الوسطاء لنفس النموذج بلوغها. وهذه بالضبط القطعة التي تجعل حلقة وكيل GEMNEX SIGHT ممكنة في نطاق أقل من ثانية للخطوة.I actually tried Cerebras — wafer-scale hardware far larger than today's GPUs, making generation vastly faster. The measured result: about 3000 tokens/second for OpenAI's 120B-sized model — a speed even brokers of the same model struggle to reach. This is exactly the piece that makes the GEMNEX SIGHT agent loop possible in the sub-second-per-step range.

SYS-08استرجاع · بحث وتطويرRetrieval · R&D🏅 تجارب متقدمةadvanced experiments
08

🔎 استرجاع معزّز (RAG) بثلاث طبقات — يفهم السؤال بثلاث عدسات معاً🔎 Three-layer retrieval-augmented generation — understanding the query through three lenses at once

نظام RAG متقدم يدمج ثلاث طبقات استرجاع — بحث بالكلمات، كلمات مفتاحية، وتشابه دلالي — ويذهب أبعد بتجارب حقيقية: تعلّم معزّز لرفع الدقة، ومسار خاص للكود عبر AST قبل التضمين، ومقارنة ممنهجة لأقوى نماذج التضمين.An advanced RAG system fusing three retrieval layers — full-text search, keywords, and semantic similarity — and going further with real experiments: reinforcement learning to raise accuracy, a dedicated code path via AST before embedding, and a systematic comparison of the strongest embedding models.

🧰 التقنياتStackHybrid RAGASTQwen3BM25RL re-rank
⚠️ المشكلةThe problem

الاسترجاع بطبقة واحدة يسقط كثيراً: التشابه الدلالي وحده يضيّع التطابقات الحرفية (أسماء دوال، رموز)، والكلمات المفتاحية وحدها تفوّت المعنى. وفي الكود تحديداً، تقطيع النص الساذج يكسر البنية فيضعف الاسترجاع — والعربية أصعب لأن أغلب النماذج مُحاباة للإنجليزية.Single-layer retrieval fails often: semantic similarity alone misses literal matches (function names, symbols), and keywords alone miss the meaning. In code specifically, naive text chunking breaks structure and weakens retrieval — and Arabic is harder because most models are biased toward English.

الحلThe solution

دمج ثلاث طبقات — كلمات بحثية + كلمات مفتاحية + تشابه دلالي — يغطّي كل طبقة نقاط ضعف الأخرى. وللكود: تحليل AST أولاً ثم تحويله إلى embeddings يحفظ البنية ويرفع جودة الاسترجاع. النتيجة استرجاع أدقّ وأكثر متانة عبر النص والشيفرة.Fusing three layers — search terms + keywords + semantic similarity — with each layer covering the others' weaknesses. And for code: AST parsing first, then converting to embeddings, preserves structure and raises retrieval quality. The result is more accurate, more robust retrieval across text and code.

🛠️ التجارب — ما جرّبته فعلاً وما تعلّمتهThe experiments — what I actually tried and learned

هذا مشروع تجارب بقدر ما هو نظام. جرّبت إضافة تعلّم معزّز (RL) فوق الاسترجاع — رفع الدقة فعلاً، لكنه أبطأ التنفيذ كثيراً، فبقي مقايضة واعية بين الدقة والسرعة. وقارنت أكثر من نموذج تضمين — E5 وQwen3 وBM25 ونماذج أخرى مخصّصة للبرمجة — فكان Qwen3 الأفضل إجمالاً. وملاحظة مهمة وثّقتها: أغلب النماذج مُحاباة للإنجليزية، فتتفوّق مؤشراتها بالإنجليزية بأكثر من 10 نقاط مئوية على العربية — قِستُ ذلك عبر مهام استرجاع معلومة مُعدّة للاختبار فقط.This is an experiments project as much as a system. I tried adding reinforcement learning (RL) on top of retrieval — it did raise accuracy, but slowed execution a lot, so it stayed a conscious accuracy-vs-speed trade-off. I compared several embedding models — E5, Qwen3, BM25 and other code-specialized models — and Qwen3 was best overall. And an important documented finding: most models are English-biased, scoring more than 10 percentage points higher in English than Arabic — I measured this via information-retrieval tasks built purely for testing.

🧩 تحت الغطاءUnder the hood
ثلاث طبقات استرجاعThree retrieval layers
بحث بالكلمات + كلمات مفتاحية + تشابه دلالي، تُدمج نتائجها لتغطية شاملة — تُمسك الحرفي والمعنى معاً.Full-text + keyword + semantic search, fused for comprehensive coverage — catching the literal and the meaning together.
مسار الكود عبر ASTCode path via AST
تحليل الشيفرة إلى شجرة AST أولاً ثم تحويلها embeddings — يحفظ البنية بدل تقطيع نصّي ساذج يكسرها.Parsing code into an AST first, then converting to embeddings — preserving structure instead of naive text chunking that breaks it.
تعلّم معزّز (RL) لرفع الدقةRL to raise accuracy
تجربة متقدمة رفعت الدقة فعلياً — بمقايضة سرعة واعية، وقرار هندسي مبني على قياس لا على حدس.An advanced experiment that genuinely raised accuracy — with a conscious speed trade-off, an engineering decision built on measurement, not intuition.
مقارنة نماذج التضمينEmbedding-model comparison
E5 · Qwen3 · BM25 وغيرها، مع نماذج مخصّصة للبرمجة — Qwen3 الأفضل إجمالاً بعد قياس ممنهج.E5 · Qwen3 · BM25 and others, plus code-specialized models — Qwen3 best overall after systematic measurement.
فجوة العربية مقابل الإنجليزيةThe Arabic-vs-English gap
قياس موثّق: تتفوّق الإنجليزية بأكثر من 10 نقاط %، ووعي مبكّر بتحدّي العربية يوجّه اختيار النموذج.A documented measurement: English leads by 10+ %, and early awareness of the Arabic challenge guides model choice.
اختبار على مهام استرجاعTesting on retrieval tasks
تقييم على مهام استرجاع معلومة مُعدّة مسبقاً لعزل جودة الاسترجاع وحدها وقياسها بموضوعية.Evaluation on pre-built information-retrieval tasks to isolate retrieval quality alone and measure it objectively.
والمزيدMore

ومشاريع أخرى قيد التطويرAnd more projects in development

ما تراه أعلاه هو محور تركيزي الحالي والأنضج تجارياً — وهناك أعمال إضافية على الطريق. تواصل معي لمعرفة المزيد أو لترتيب عرض مباشر. 🤝What you see above is my current focus and the most commercially mature — with more work on the way. Reach out to learn more or to arrange a live demo. 🤝