منصة توسيم وتدريب رؤية حاسوبية تعمل بالكامل داخل شبكة العميل، تغطي دورة الحياة الكاملة: جمع البيانات ← التوسيم الآلي ← التحقق ← التدريب ← التقييم ← إعادة التوسيم بالنموذج المُدرَّب. لا تُخرج صورة أو توسيماً خارج الجهاز — الخصوصية خاصية معمارية قابلة للتدقيق، لا بند في سياسة.A computer-vision labeling & training platform that runs entirely inside the client's network, covering the full lifecycle: data collection → auto-labeling → verification → training → evaluation → re-labeling with the trained model. No image or annotation ever leaves the machine — privacy is an auditable architectural property, not a policy clause.
🎬 شاهد ديمو NUWA Labeling🎬 Watch the NUWA Labeling demoشركات الرؤية الحاسوبية تعمل على لقطات CCTV وبيانات حسّاسة، ورفعها لمنصة سحابية عائق تعاقدي وتنظيمي — لا مجرد تفضيل. ثم يأتي ألم البداية من الصفر (cold start): توسيم يدوي يكلّف أياماً قبل أن يرى النموذج أول صورة، وتكلفة لكل مقعد ولكل صورة تجعل المشروع الكبير غير مجدٍ أصلاً.Computer-vision teams work on CCTV footage and sensitive data, and uploading it to a cloud platform is a contractual and regulatory blocker — not just a preference. Then comes the cold-start pain: manual labeling costs days before the model sees its first image, and per-seat, per-image pricing makes large projects unviable to begin with.
منصة on-premises تغطي الدورة كاملة محلياً: التوسيم الآلي يبدأ بلا صورة موسومة واحدة عبر Grounding DINO، فيتقلّص دور الإنسان من الرسم إلى الموافقة. التكلفة الحدّية بعد التركيب تقارب الصفر، وعلى المجموعات الكبيرة يتحوّل هذا من توفير هامشي إلى فارق في جدوى المشروع نفسه.An on-premises platform covering the full cycle locally: auto-labeling starts with zero labeled images via Grounding DINO, shrinking the human's role from drawing to approving. Marginal cost after install is near zero — and on large datasets this shifts from marginal savings to the difference in whether the project is viable at all.
المشكلة الجوهرية في أدوات التوسيم المكتبية: استدلال PyTorch يمسك الـ GIL فيتجمّد التطبيق أثناء العمل الثقيل. الحل المطبَّق هنا يفصل ثلاث طبقات تزامن — واجهة، خيوط إشراف، وعمليات استدلال معزولة بذاكرتها وسياق CUDA خاص بها — فتبقى الواجهة حيّة تحت أثقل حمل. هذه الصلابة الهندسية هي ما حوّلها من نموذج أوّلي إلى منصة إنتاجية (~25 ألف سطر) بُني عليها 10 منتجات تُستخدم تجارياً وتدرّ دخلاً.The core problem in desktop labeling tools: PyTorch inference holds the GIL, freezing the app during heavy work. The solution here separates three concurrency layers — UI, supervisor threads, and isolated inference processes each with their own memory and CUDA context — so the interface stays alive under the heaviest load. That engineering robustness is what turned it from a prototype into a production platform (~25K lines) on top of which 10 commercially-used, revenue-generating products were built.
محاوِرة افتراضية ثلاثية الأبعاد تتكلم بلهجة سعودية، تدير المقابلة الأولى كاملة لكل مرشّح بجودة ثابتة، وتُسلّم تقرير تقييم فور انتهائها. النتيجة: فريق التوظيف يبدأ من قائمة مختصرة مسنودة بحوار فعلي — لا بقراءة ست ثوانٍ لسيرة ذاتية.A 3D virtual interviewer that speaks in a Saudi dialect, runs the full first-round interview for every candidate at consistent quality, and delivers an assessment report the moment it ends. The result: the hiring team starts from a shortlist backed by an actual conversation — not a six-second résumé skim.
🎬 شاهد ديمو «زينة» الكامل🎬 Watch the full "Zeina" demoمئات السِيَر تدخل، وطاقة الفريق البشري لا تتجاوز عدداً محدوداً من المقابلات. النتيجة ثلاث خسائر في آنٍ واحد: مرشّحون أكفاء يُستبعدون لأن أحداً لم يفتح ملفهم، وأسابيع تُهدر في مقابلات أولى متكررة على نفس الأسئلة، ومرشّح ينتظر بلا رد فتخسر الشركة سمعتها كجهة توظيف.Hundreds of résumés come in, and a human team can only run a limited number of interviews. The result is three losses at once: strong candidates dropped because no one opened their file, weeks wasted on repetitive first-round interviews asking the same questions, and candidates left waiting — costing the company its employer reputation.
مقابلة أولى آلية لكل مرشّح بلا استثناء وبنفس الجودة: حوار طبيعي، أسئلة مبنية على سيرته والوظيفة، وتقرير تقييم فوري. يتحوّل عمل الفريق البشري من غربلة مئات الملفات إلى اتخاذ قرار على قائمة قصيرة موثوقة — أسابيع تنكمش إلى ساعات.An automated first-round interview for every candidate, without exception, at the same quality: a natural conversation, questions grounded in their CV and the role, and an instant assessment report. The human team's job shifts from screening hundreds of files to deciding over a trusted shortlist — weeks shrink to hours.
مقابلة حقيقية لا تحتمل تأخيراً محرجاً ولا هلوسة تُهين المرشّح. فبنيت كل طبقة لخدمة إحساس واحد: أنك تحاور إنساناً. لذلك نموذجا Whisper يعملان بالتوازي حتى لا يُمضغ لهجتك أو مصطلحك، والبث عبر WebSocket بـ ~150ms حتى لا تنتظر ردّها، والاستدلال على Groq LPU حتى يفكّر السؤال التالي أسرع مما تُكمل جوابك. والأهم — طبقة عدالة تضمن أن يُقيَّم كل مرشّح على ما قاله فعلاً، لا على خطأ في تحويل الصوت. جرّبها متطوّعون حقيقيون، وجرّبتها بنفسي: التجربة ممتعة والمستوى فخم فعلاً.A real interview can't tolerate an awkward delay or a hallucination that insults the candidate. So every layer was built to serve one feeling: that you're talking to a human. Two Whisper models run in parallel so your dialect or technical term isn't garbled; streaming over WebSocket at ~150ms so you don't wait for its reply; inference on Groq LPU so the next question is thought up faster than you finish answering. Most important — a fairness layer ensures each candidate is judged on what they actually said, not on a speech-to-text error. Real volunteers tried it, and so did I: the experience is genuinely enjoyable and the quality feels premium.
تفتح الرابط — بلا تحميل تطبيق ولا تسجيلOpen the link — no app install, no signup
زينة ترحّب بك بصوت طبيعي وتبدأ الحوارZeina greets you in a natural voice and starts the conversation
كلامك يتحول لنص وأنت تتكلم — لحظياًYour speech turns to text as you talk — in real time
أسئلة من سيرتك أنت والوظيفة، تتعمّق مع كل إجابةQuestions drawn from your CV and the role, deepening with each answer
أجندة حية تريك ما غُطّي وما تبقّى — تستدرك بنفسكA live agenda shows what's covered and what's left — it self-corrects
تقرير تقييمك كاملاً لحظة انتهاء المقابلةYour full assessment report the moment the interview ends
نظام رؤية حاسوبية يبني توأماً رقمياً للبيئة لحظياً: يكتشف كل شخص داخل منطقة محددة، يحوّله إلى شخصية 3D تتبع حركته الفعلية، يتعرّف عليه بوجهه عبر كاميرات CCTV متعددة، ويعرض فوق شخصيته اسمه ورتبته وصورته الملتقطة آنياً.A computer-vision system that builds a real-time digital twin of the environment: it detects every person in a defined area, turns them into a 3D character that follows their actual movement, recognizes them by face across multiple CCTV cameras, and displays their name, rank and a live-captured photo above their character.
كاميرات CCTV التقليدية تسجيل سلبي بلا فهم: لا تعرف مَن الشخص، تفقد أثره حين يستدير أو يمرّ خلف عائق، ولا تربط اللقطات بين كاميرات متعددة. النتيجة مراقبة بلا وعي لحظي بمن هو موجود وأين، وحضور يُرصد يدوياً.Traditional CCTV is passive recording with no understanding: it doesn't know who the person is, loses track when they turn or pass behind an obstacle, and never links footage across cameras. The result is surveillance with no live awareness of who is present and where, and attendance tracked by hand.
توأم رقمي 3D يعيد بناء المشهد حياً: كل شخص شخصية تتحرك بحركته، معرّفة بالاسم والرتبة وصورة آنية فوق رأسه، ومتتبَّعة عبر عدة كاميرات دون فقدان الأثر. أُثبت عملياً على أكثر من 10 أشخاص في أوقات مختلفة بأداء ممتاز — ويصلح كذلك كنظام حضور تلقائي فعّال.A 3D digital twin that rebuilds the scene live: each person is a character that moves as they move, tagged with name, rank and a live photo above their head, and tracked across multiple cameras without losing the trail. Proven in practice on 10+ people at different times with excellent performance — and it also works as an effective automatic attendance system.
أصعب تحدٍّ في التتبع هو فقدان الوجه: يختفي حين يستدير الشخص أو يبتعد فينكسر التعرّف. لذلك بنيت مساراً هجيناً متيناً — InsightFace يتعرّف بالوجه أولاً، فإذا فُقد الوجه يواصل التتبع عبر صندوق الجسم من YOLO nano (26) فلا ينقطع الأثر. وللتعرّف القوي بدل صورة واحدة: عدة لقطات لكل شخص + augmentation، تُخزَّن كـ embeddings وتُسترجع عبر FAISS بسرعة ودقة عالية حتى مع اختلاف الإضاءة والزاوية. هذه التركيبة تحديداً هي ما جعل الأداء ممتازاً عبر كاميرات وأوقات مختلفة.The hardest tracking challenge is losing the face: it disappears when the person turns or moves away, breaking recognition. So I built a robust hybrid path — InsightFace recognizes the face first, and if the face is lost it keeps tracking via the body box from YOLO nano (26) so the trail never breaks. For robust recognition, instead of a single image: several shots per person + augmentation, stored as embeddings and retrieved via FAISS quickly and accurately even under different lighting and angles. That exact combination is what made performance excellent across cameras and times.
جرّبت عدة إيماءات: إشارة «OK» تعطي «تم»، لكنها أضعف على المسافات البعيدة. أما رفع اليد فوق الرأس فكان ممتازاً حتى من بعيد، فاعتمدته لضبط حالة «مشغول» مع إشعار يُسجَّل في الواجهة (تعمل على شاشة المكتب). ترتفع الحالة تلقائياً بعد نحو ساعة أو بتكرار الإيماءة. وأثبت النظام فاعليته أيضاً في بيئة الحضور التلقائي.I tried several gestures: an "OK" sign for "done", but it was weaker at long range. Raising a hand above the head worked excellently even from afar, so I adopted it to set "busy" status with a notification logged in the UI (runs on a desktop screen). The status auto-clears after about an hour or by repeating the gesture. The system also proved effective in an automatic attendance setting.
المنتج ليس OCR — بل وكيل يرى الشاشة، يفهمها، ويتحكّم بها. هذا المحرّك هو العين: طبقة إدراك (OCR + كشف عناصر الواجهة) مدرَّبة بالكامل من الصفر، تعمل محلياً في ~0.4 ثانية، وتغذّي دماغاً لغوياً على Cerebras — فتصير الحلقة الكاملة في نطاق أقل من ثانية للخطوة.The product isn't OCR — it's an agent that sees the screen, understands it, and controls it. This engine is the eye: a perception layer (OCR + UI-element detection) trained entirely from scratch, running locally in ~0.4s, feeding a language brain on Cerebras — so the full loop lands in the range of under one second per step.
وكلاء الواجهات اليوم بطيئون، والبطء من طرفين: الإدراك (محرّكات OCR عامة أو نماذج VLM ضخمة تأخذ ثوانٍ لكل لقطة) والتخطيط (نماذج لغوية على بنية تقليدية). والحلقة تسير بسرعة أبطأ مكوّن فيها — فتسريع طرف واحد لا يعطي شيئاً، والنتيجة وكيل «انتظر ثم شاهد» لا وكيل يعمل.GUI agents today are slow, and the slowness comes from two ends: perception (general OCR engines or huge VLMs taking seconds per frame) and planning (language models on conventional hardware). The loop runs at the speed of its slowest component — so speeding up one end gives you nothing, and you get a "wait-then-watch" agent, not an agent that works.
عالجنا الطرفين معاً: إدراك محلّي مخصّص في ~0.4 ثانية + تخطيط فائق السرعة على Cerebras. النتيجة انتقال من إيقاع «انتظر ثم شاهد» إلى تفاعل آنيّ حقيقي — ووضع خصوصية صارم: العين لا ترسل بكسلاً واحداً خارج الجهاز، والنموذج اللغوي يستقبل نصاً وإحداثيات فقط لا صورة الشاشة.We addressed both ends together: custom local perception in ~0.4s + ultra-fast planning on Cerebras. The result is a shift from "wait-then-watch" to true real-time interaction — with a strict privacy mode: the eye sends not a single pixel off-device, and the language model receives only text and coordinates, never the screen image.
جرّبت المحرّكات العالمية أولاً — EasyOCR وPaddleOCR وRapidOCR — فوجدتها تأخذ ثوانيَ على اللقطة الواحدة، وهو ما يجعل بناء وكيل تفاعلي فوقها مستحيلاً عملياً. لقطة الواجهة نوع مختلف كليّاً عن صور المستندات: نصّ صغير جداً، كثافة مئات الصناديق، أيقونات ذات معنى، وعربية وإنجليزية تختلطان في السطر. فقرّرت التدريب من الصفر: كاشف YOLO26 مضبوط على كثافة الواجهات، متعرّف CRNN+CTC بمعالجة عربية دقيقة، وخطّ استدلال بأشكال ثابتة — حتى نزل الزمن من 5 ثوانٍ إلى ~0.4. كل ذلك لسبب واحد: الوصول إلى وكيل سريع جداً وبأداء قوي.I tried the mainstream engines first — EasyOCR, PaddleOCR, RapidOCR — and found they take seconds per frame, which makes building an interactive agent on top of them practically impossible. A UI screenshot is a completely different kind of image from document scans: tiny text, hundreds of dense boxes, meaningful icons, and Arabic and English mixed in one line. So I decided to train from scratch: a YOLO26 detector tuned for UI density, a CRNN+CTC recognizer with careful Arabic processing, and a fixed-shape inference pipeline — until latency dropped from 5 seconds to ~0.4. All for one reason: reaching a very fast, strong agent.
📐 كيف تُقرأ: قياس بعد تسخين كل محرّك، والرقم وسيط عدّة تشغيلات. جزء من الفارق أن نظامنا على GPU؛ لكن الفارق البنيوي الحقيقي من التخصّص: كاشف مضبوط على كثافة الواجهات، متعرّف بعرض قصير ملائم لأسطرها، وخطّ استدلال بأشكال ثابتة مسخّنة مسبقاً. عند ~0.4 ثانية للخطوة يصير التفاعل الآني ممكناً — وعند عدّة ثوانٍ لا يصير. هذا فارق في رتبة الحجم.📐 How to read it: measured after warming up each engine, with the median of several runs. Part of the gap is that our system is on GPU; but the real structural gap is from specialization: a detector tuned for UI density, a recognizer with a short width suited to UI lines, and a fixed-shape, pre-warmed inference pipeline. At ~0.4s per step, real-time interaction becomes possible — at several seconds, it doesn't. This is an order-of-magnitude difference.
تطبيق مكتبي يعمل بالكامل على جهاز العميل: يحوّل ملف بيانات خام (CSV / Excel / JSON) إلى لوحة تحليل بصرية + نموذج تعلّم آلي مدرَّب + توقّعات مستقبلية + ملف نموذج جاهز للنشر — دون سطر برمجي واحد، ودون أن يغادر بايت واحد جهازك.A desktop app that runs entirely on the client's machine: it turns a raw data file (CSV / Excel / JSON) into a visual analytics dashboard + a trained ML model + future forecasts + a deployment-ready model file — without a single line of code, and without a single byte leaving your machine.
🎬 شاهد ديمو GEMNEX AI/ML🎬 Watch the GEMNEX AI/ML demoكل تحليل يتطلّب مختص Python/BI متفرّغاً، وكل سؤال بسيط يتحوّل إلى دورة طلب ← انتظار ← تسليم تمتد أياماً. أدوات السحابة مرفوضة أصلاً في القطاعات الحسّاسة، والأخطر: تُعرض أرقام دقة مرتفعة بلا تدقيق، فتُبنى قرارات على نماذج مضلِّلة تنهار في الإنتاج.Every analysis needs a dedicated Python/BI specialist, and every simple question turns into a request → wait → deliver cycle lasting days. Cloud tools are outright rejected in sensitive sectors, and worse: high accuracy numbers are shown without scrutiny, so decisions get built on misleading models that collapse in production.
«ارفع ملفك، وخلال دقائق تحصل على المسار الكامل»: تنظيف تلقائي، لوحة رسوم فورية، عدّة نماذج مدرَّبة ومقارنة، توقّعات مستقبلية، وملف نموذج جاهز للنشر — بلا كود، وبلا سحابة. والاعتماد على مختص متفرّغ يتحوّل من إلزامي إلى اختياري."Upload your file, and within minutes you get the full pipeline": automatic cleaning, an instant charts dashboard, several trained and compared models, future forecasts, and a deployment-ready model file — no code, no cloud. Relying on a dedicated specialist shifts from mandatory to optional.
الخوارزميات ليست الفجوة — فهي متاحة ومجانية. الفجوة في المسار الكامل من الملف الخام إلى قرار موثوق، وفي الثقة بأن الرقم المعروض صادق. بُني على ~3 أشهر و14 دورة تحسين موثّقة، كل واحدة انطلقت من استخدام حقيقي كشف مشكلة حقيقية، وكل إصلاح وثّق السبب الجذري لا العَرَض. أي أداة تُخرج رقم دقة — القليل منها يخبرك متى يكون هذا الرقم كذبة، وهذه بالضبط ميزتها التنافسية.The algorithms aren't the gap — they're free and available. The gap is the full pipeline from raw file to a trustworthy decision, and the confidence that the number shown is honest. Built over ~3 months and 14 documented improvement cycles, each starting from real usage that exposed a real problem, and each fix documenting the root cause, not the symptom. Any tool outputs an accuracy number — few tell you when that number is a lie, and that is exactly its competitive edge.
خط إنتاج آلي يحوّل الصوت الخام إلى بيانات تدريب ASR عالية الجودة: ثلاثة محكّمين من أقوى نماذج التعرّف على الكلام، تحديد عدد المتحدثين، وتوقيت كل كلمة — ثم مراجعة وتعديل بمزوّدين، ولا يمرّ للتدريب إلا ما تجاوزت دقته 70%.An automated pipeline that turns raw audio into high-quality ASR training data: three judges from the strongest speech-recognition models, speaker counting, and per-word timing — then review and correction by two providers, and only data above 70% accuracy passes to training.
سوق نماذج تحويل الكلام إلى نص عليه طلب عالٍ — خصوصاً بالعربية — لكن البيانات عالية الجودة نادرة ومحدودة. إنتاجها يدوياً يحتاج فرق توسيم كاملة، ووقتاً طويلاً، وتكلفة مرتفعة تجعل بناء نموذج عربي جيد عائقاً بذاته.The speech-to-text market has high demand — especially in Arabic — but high-quality data is scarce and limited. Producing it by hand needs entire labeling teams, long timelines, and high costs that make building a good Arabic model a blocker in itself.
خط إنتاج يحوّل الصوت الخام إلى بيانات تدريب موثوقة بأقل تدخّل بشري — يرفع الجودة، يقلّص الاعتماد على الموظفين، ويملأ فجوة السوق العربي بنسبة عالية. المخرَج بيانات جاهزة للتدريب، لا مجرد تفريغ صوتي.A pipeline that turns raw audio into reliable training data with minimal human intervention — raising quality, reducing reliance on staff, and largely filling the Arabic market gap. The output is training-ready data, not just a transcript.
لا أعتمد على نموذج واحد قد يخطئ: ثلاثة محكّمين من أقوى نماذج ASR يحكّمون على النص معاً، فتُلتقط الكلمة الصحيحة حتى حين يتعثّر أحدها. طبقة تحديد المتحدثين من مكتبة متخصصة على HuggingFace تفصل من قال ماذا، وWhisperX يستخرج توقيت كل كلمة على حدة لا الجملة. ثم مزوّدان قويان: أحدهما يراجع والآخر يعدّل، ولا يُعتمَد للتدريب إلا ما تجاوزت دقته 70% — مع خيار مراجعة يدوية للأجزاء الأعقد. كل ذلك يعمل على خدمة سحابية قابلة للتوسّع.I don't rely on a single model that might err: three judges from the strongest ASR models arbitrate the text together, so the correct word is captured even when one stumbles. A speaker-diarization layer from a specialized HuggingFace library separates who said what, and WhisperX extracts per-word — not per-sentence — timing. Then two strong providers: one reviews and the other corrects, and only data above 70% accuracy is accepted for training — with a manual-review option for the trickiest parts. All of it runs on a scalable cloud service.
مشروع بحثي بنيته من الصفر عبر عشرات التجارب على عتاد سحابي عملاق: نموذج Transformers بانتباه كامل وترميز BPE مخصّص للعربية والإنجليزية وبايثون، وصل نتائج جيدة نسبةً لحجمه على نطاق محدّد (sub-domain). ليس استدعاء API — بل بناءٌ يثبت فهم البنية بالكامل.An R&D project I built from scratch through dozens of experiments on massive cloud hardware: a full-attention Transformer with a custom BPE tokenizer for Arabic, English and Python, reaching good results relative to its size on a specific sub-domain. Not an API call — a build that proves complete understanding of the architecture.
بناء نموذج لغوي كفؤ من الصفر — لا استدعاء API جاهز — على ميزانية عتاد محدودة تفرض تجارب سريعة ومركّزة، مع توكنايزر يخدم العربية والإنجليزية والبرمجة معاً في آنٍ واحد.Building a competent language model from scratch — not a ready API call — on a limited hardware budget that forces fast, focused experiments, with a tokenizer serving Arabic, English and code all at once.
نموذج Transformers بانتباه كامل وترميز BPE مخصّص، وصل نقطة جيدة نسبةً لحجمه على نطاق محدّد بعد سلسلة تجارب ممنهجة بمؤشرات أداء وطباعة عيّنات — وخبرة عملية عميقة بآليات الانتباه وتقنيات التسريع الحديثة.A full-attention Transformer with a custom BPE tokenizer, reaching a good point relative to its size on a specific domain after a systematic series of experiments with performance metrics and sample printouts — plus deep hands-on experience with attention mechanisms and modern acceleration techniques.
جرّبت Cerebras فعلياً — عتاد بمقياس الرقاقة (wafer-scale) أكبر بكثير من كروت الشاشة الحالية، فيصير التوليد أسرع بمراحل. النتيجة المقيسة: نحو 3000 توكن/ثانية لنموذج بحجم 120B من OpenAI — سرعة يصعب حتى على الوسطاء لنفس النموذج بلوغها. وهذه بالضبط القطعة التي تجعل حلقة وكيل GEMNEX SIGHT ممكنة في نطاق أقل من ثانية للخطوة.I actually tried Cerebras — wafer-scale hardware far larger than today's GPUs, making generation vastly faster. The measured result: about 3000 tokens/second for OpenAI's 120B-sized model — a speed even brokers of the same model struggle to reach. This is exactly the piece that makes the GEMNEX SIGHT agent loop possible in the sub-second-per-step range.
نظام RAG متقدم يدمج ثلاث طبقات استرجاع — بحث بالكلمات، كلمات مفتاحية، وتشابه دلالي — ويذهب أبعد بتجارب حقيقية: تعلّم معزّز لرفع الدقة، ومسار خاص للكود عبر AST قبل التضمين، ومقارنة ممنهجة لأقوى نماذج التضمين.An advanced RAG system fusing three retrieval layers — full-text search, keywords, and semantic similarity — and going further with real experiments: reinforcement learning to raise accuracy, a dedicated code path via AST before embedding, and a systematic comparison of the strongest embedding models.
الاسترجاع بطبقة واحدة يسقط كثيراً: التشابه الدلالي وحده يضيّع التطابقات الحرفية (أسماء دوال، رموز)، والكلمات المفتاحية وحدها تفوّت المعنى. وفي الكود تحديداً، تقطيع النص الساذج يكسر البنية فيضعف الاسترجاع — والعربية أصعب لأن أغلب النماذج مُحاباة للإنجليزية.Single-layer retrieval fails often: semantic similarity alone misses literal matches (function names, symbols), and keywords alone miss the meaning. In code specifically, naive text chunking breaks structure and weakens retrieval — and Arabic is harder because most models are biased toward English.
دمج ثلاث طبقات — كلمات بحثية + كلمات مفتاحية + تشابه دلالي — يغطّي كل طبقة نقاط ضعف الأخرى. وللكود: تحليل AST أولاً ثم تحويله إلى embeddings يحفظ البنية ويرفع جودة الاسترجاع. النتيجة استرجاع أدقّ وأكثر متانة عبر النص والشيفرة.Fusing three layers — search terms + keywords + semantic similarity — with each layer covering the others' weaknesses. And for code: AST parsing first, then converting to embeddings, preserves structure and raises retrieval quality. The result is more accurate, more robust retrieval across text and code.
هذا مشروع تجارب بقدر ما هو نظام. جرّبت إضافة تعلّم معزّز (RL) فوق الاسترجاع — رفع الدقة فعلاً، لكنه أبطأ التنفيذ كثيراً، فبقي مقايضة واعية بين الدقة والسرعة. وقارنت أكثر من نموذج تضمين — E5 وQwen3 وBM25 ونماذج أخرى مخصّصة للبرمجة — فكان Qwen3 الأفضل إجمالاً. وملاحظة مهمة وثّقتها: أغلب النماذج مُحاباة للإنجليزية، فتتفوّق مؤشراتها بالإنجليزية بأكثر من 10 نقاط مئوية على العربية — قِستُ ذلك عبر مهام استرجاع معلومة مُعدّة للاختبار فقط.This is an experiments project as much as a system. I tried adding reinforcement learning (RL) on top of retrieval — it did raise accuracy, but slowed execution a lot, so it stayed a conscious accuracy-vs-speed trade-off. I compared several embedding models — E5, Qwen3, BM25 and other code-specialized models — and Qwen3 was best overall. And an important documented finding: most models are English-biased, scoring more than 10 percentage points higher in English than Arabic — I measured this via information-retrieval tasks built purely for testing.
ما تراه أعلاه هو محور تركيزي الحالي والأنضج تجارياً — وهناك أعمال إضافية على الطريق. تواصل معي لمعرفة المزيد أو لترتيب عرض مباشر. 🤝What you see above is my current focus and the most commercially mature — with more work on the way. Reach out to learn more or to arrange a live demo. 🤝