8000
Skip to content
View avijit9's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report avijit9

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

LAVIS - A One-stop Library for Language-Vision Intelligence

Jupyter Notebook 11,262 1,108 Updated Jun 2, 2026

🎥 [Awesome] Egocentric / First-Person Video Datasets 📚 Papers, Benchmarks & Resources for Ego Vision

209 6 Updated Aug 22, 2026

Convert 2D videos and photos into interactive 3D scenes using Gaussian splatting backends (VGGT, LongSplat, DepthSplat, SHARP, TripoSplat), Rerun, and the SuperSplat editor.

Python 154 9 Updated Jul 9, 2026

Code for "EgoX: Egocentric Video Generation from a Single Exocentric Video"

Python 751 51 Updated Jul 10, 2026

This is a collection of recent papers on reasoning in video generation models.

164 6 Updated Aug 24, 2026

DuoLoRA implementation

Python 9 Updated Oct 18, 2025

pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras. It provides a broad set of modern local and global feature extractors, multiple loop-closure stra…

Python 3,401 542 Updated Aug 23, 2026

Official PyTorch implementation of the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs"

Python 100 15 Updated Jun 6, 2025

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 26,237 2,055 Updated Aug 25, 2026

[CVPR 2025] Official PyTorch code of "Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation".

Python 59 Updated Jul 26, 2026

[ICCV 2023] UniVTG: Towards Unified Video-Language Temporal Grounding

Python 379 32 Updated May 8, 2024

Group-wise Temporal Logit Adjustment for TAS

Python 12 Updated Oct 24, 2024

A curated list for awesome discrete diffusion models resources.

573 28 Updated Sep 9, 2025

[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer

Python 14,287 1,537 Updated May 19, 2026

Building simple diffusion models for image generation. More so for understanding and learning.

Python 8 2 Updated Mar 30, 2025

[WACV'25] Temporal Instructional Diagram Grounding in Unconstrained Videos

Python 5 Updated Dec 17, 2024

Video Annotation Tool

Vue 242 34 Updated Jul 27, 2026

[ICLR 2025] Video Action Differencing

Python 53 6 Updated Jul 3, 2025

A collection of my book notes on various subjects, mainly computer science

Java 3,055 805 Updated Jun 13, 2026

[ECCV2024] Gated Temporal Action Anticipation for Stochastic Long-Term Anticipation

Python 24 1 Updated May 29, 2025

Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment" (NeurIPS 2023)

Python 19 4 Updated Apr 5, 2024

React + Next.js template for research websites (for PhD students, researchers, etc)

TypeScript 236 101 Updated Jan 12, 2025

Visualizing the learned space-time attention using Attention Rollout

Jupyter Notebook 42 7 Updated Apr 1, 2022

MLX: An array framework for Apple silicon

C++ 28,158 2,182 Updated Aug 26, 2026
Jupyter Notebook 217 21 Updated Nov 10, 2024

Collection of AWESOME vision-language models for vision tasks

3,128 234 Updated Oct 14, 2025

A declarative drawing API in Python

Python 301 14 Updated Aug 28, 2024

It is my belief that you, the postgraduate students and job-seekers for whom the book is primarily meant will benefit from reading it; however, it is my hope that even the most experienced research…

4,890 325 Updated Aug 22, 2025
Next
0