在使用我们的模型之前,您需要先确保环境中已安装所有必要的依赖项。这些依赖项涵盖了模型运行所需的各类库和工具,确保您可以顺利进行模型推理。
请按照以下步骤进行安装:
pip install -r requirements.txt
安装完所有必要的依赖项后,您就可以开始使用我们的模型进行推理了。我们提供了两种推理方式:使用终端进行推理和使用交互式推理。
这里我们以示例图片asserts/demo.jpg为例进行说明:
如果您希望直接在终端中运行推理脚本,可以使用以下命令:
python chatme.py --image asserts/demo.jpg --question "货架上有几个苹果?"
此命令会加载预训练的模型,并使用提供的图片(demo.jpg)和问题("货架上有几个苹果?")进行推理。
模型会分析图片并尝试回答提出的问题,推理结果将以文本形式输出到终端中,例如:
小千:货架上有三个苹果。
除了使用终端进行推理,您还可以使用交互式推理功能与大模型进行实时交互。要启动交互式终端,请运行以下命令:
python main.py
此命令会启动一个交互式终端,等待您输入图片地址。您可以在终端中输入图片地址(例如asserts/demo.jpg),然后按下回车键。
模型会根据您提供的图片进行推理,并等待您输入问题。
一旦您输入了问题(例如"货架上有几个苹果?"),模型就会分析图片并尝试回答,推理结果将以文本形式输出到终端中,例如:
图片地址 >>>>> asserts/demo.jpg
用户:货架上有几个苹果?
小千:货架上有三个苹果。
通过这种方式,您可以轻松地与模型进行交互,并向其提出各种问题。
Towards Real-World Test-Time Adaptation: Tri-Net Self-Training with Balanced Normalization
Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored Clustering
Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection
Intra- and Inter-Slice Contrastive Learning for Point Supervised OCT Fluid Segmentation
Partitioning Stateful Data Stream Applications in Dynamic Edge Cloud Environments
Closed-loop Matters: Dual Regression Networks for Single Image Super-Resolution
Graph Convolutional Networks for Temporal Action Localization
NAT: Neural Architecture Transformer for Accurate and Compact Architectures
Breaking the Curse of Space Explosion: Towards Effcient NAS with Curriculum Search
Contrastive Neural Architecture Search with Neural Architecture Comparators
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation
Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification
Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation Score
Masked Motion Encoding for Self-Supervised Video Representation Learning
Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation
Prototype-Guided Continual Adaptation for Class-Incremental Unsupervised Domain Adaptation
Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection
Polysemy Deciphering Network for Human-Object Interaction Detection
Bidirectional Posture-Appearance Interaction Network for Driver Behavior Recognition
Improving Generative Adversarial Networks with Local Coordinate Coding
SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation
Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial Attention
Deep Multi-View Learning Using Neuron-Wise Correlation-Maximizing Regularizers
Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic Segmentation
Contextual Point Cloud Modeling for Weakly-supervised Point Cloud Semantic Segmentation
Quasi-Balanced Self-Training on Noise-Aware Synthesis of Object Point Clouds for Closing Domain Gap
Test-Time Model Adaptation for Visual Question Answering with Debiased Self-Supervisions
Debiased Visual Question Answering from Feature and Sample Perspectives
Intelligent Home 3D: Automatic 3D-House Design from Linguistic Descriptions Only
Cross-Modal Relation-Aware Networks for Audio-Visual Event Localization
Cascade Reasoning Network for Text-based Visual Question Answering
Python
100.0%
在使用我们的模型之前,您需要先确保环境中已安装所有必要的依赖项。这些依赖项涵盖了模型运行所需的各类库和工具,确保您可以顺利进行模型推理。
请按照以下步骤进行安装:
pip install -r requirements.txt
安装完所有必要的依赖项后,您就可以开始使用我们的模型进行推理了。我们提供了两种推理方式:使用终端进行推理和使用交互式推理。
这里我们以示例图片asserts/demo.jpg为例进行说明:
如果您希望直接在终端中运行推理脚本,可以使用以下命令:
python chatme.py --image asserts/demo.jpg --question "货架上有几个苹果?"
此命令会加载预训练的模型,并使用提供的图片(demo.jpg)和问题("货架上有几个苹果?")进行推理。
模型会分析图片并尝试回答提出的问题,推理结果将以文本形式输出到终端中,例如:
小千:货架上有三个苹果。
除了使用终端进行推理,您还可以使用交互式推理功能与大模型进行实时交互。要启动交互式终端,请运行以下命令:
python main.py
此命令会启动一个交互式终端,等待您输入图片地址。您可以在终端中输入图片地址(例如asserts/demo.jpg),然后按下回车键。
模型会根据您提供的图片进行推理,并等待您输入问题。
一旦您输入了问题(例如"货架上有几个苹果?"),模型就会分析图片并尝试回答,推理结果将以文本形式输出到终端中,例如:
图片地址 >>>>> asserts/demo.jpg
用户:货架上有几个苹果?
小千:货架上有三个苹果。
通过这种方式,您可以轻松地与模型进行交互,并向其提出各种问题。
Towards Real-World Test-Time Adaptation: Tri-Net Self-Training with Balanced Normalization
Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored Clustering
Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection
Intra- and Inter-Slice Contrastive Learning for Point Supervised OCT Fluid Segmentation
Partitioning Stateful Data Stream Applications in Dynamic Edge Cloud Environments
Closed-loop Matters: Dual Regression Networks for Single Image Super-Resolution
Graph Convolutional Networks for Temporal Action Localization
NAT: Neural Architecture Transformer for Accurate and Compact Architectures
Breaking the Curse of Space Explosion: Towards Effcient NAS with Curriculum Search
Contrastive Neural Architecture Search with Neural Architecture Comparators
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation
Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification
Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation Score
Masked Motion Encoding for Self-Supervised Video Representation Learning
Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation
Prototype-Guided Continual Adaptation for Class-Incremental Unsupervised Domain Adaptation
Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection
Polysemy Deciphering Network for Human-Object Interaction Detection
Bidirectional Posture-Appearance Interaction Network for Driver Behavior Recognition
Improving Generative Adversarial Networks with Local Coordinate Coding
SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation
Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial Attention
Deep Multi-View Learning Using Neuron-Wise Correlation-Maximizing Regularizers
Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic Segmentation
Contextual Point Cloud Modeling for Weakly-supervised Point Cloud Semantic Segmentation
Quasi-Balanced Self-Training on Noise-Aware Synthesis of Object Point Clouds for Closing Domain Gap
Test-Time Model Adaptation for Visual Question Answering with Debiased Self-Supervisions
Debiased Visual Question Answering from Feature and Sample Perspectives
Intelligent Home 3D: Automatic 3D-House Design from Linguistic Descriptions Only
Cross-Modal Relation-Aware Networks for Audio-Visual Event Localization
Cascade Reasoning Network for Text-based Visual Question Answering
Python
100.0%