NeuroHash: A Hyperdimensional Neuro-Symbolic Framework for Spatially-Aware Image
Hashing and Retrieval
Abstract
Customizable image retrieval from large datasets remains a critical challenge, particularly when preserving spatial relationships within images. Traditional hashing methods, primarily based on deep learning, often fail to capture spatial information adequately and lack transparency. In this paper, we introduce NeuroHash , a novel neuro-symbolic framework leveraging Hyperdimensional Computing (HDC) to enable highly customizable, spatially-aware image retrieval. NeuroHash combines pre-trained deep neural network models with HDC-based symbolic models, allowing for flexible manipulation of hash values to support conditional image retrieval. Our method includes a self-supervised context-aware HDC encoder and novel loss terms for optimizing lower-dimensional bipolar hashing using multilinear hyperplanes. We evaluate NeuroHash on two benchmark datasets, demonstrating superior performance compared to state-of-the-art hashing methods, as measured by mAP@5K scores and our newly introduced metric, mAP@5Kr, which assesses spatial alignment. The results highlight NeuroHash ’s ability to achieve competitive performance while offering significant advantages in flexibility and customization, paving the way for more advanced and versatile image retrieval systems.
I Introduction
In the era of explosive growth in image data, managing vast repositories of images, particularly in domains requiring the swift retrieval of similar images for a given query, presents an escalating challenge. Numerous research endeavors have sought to develop efficient and accurate methods for similar image retrieval. The primary focus of these investigations has been the design of adept hash functions capable of transforming images into a compact, fixed-size hash, thereby encapsulating their similarity to other images.
One early research utilized shallow machine learning models such as support vector machine (SVM) to extract discrete features from each image to hashing images [29]. As deep neural networks (DNNs) show remarkable performance on various image-based tasks, image hashing models for image retrieval based on neural networks are proposed starting from Convolutional Neural Networks (CNNs) based approaches [37, 2]. This momentum of applying DNNs to image retrieval tasks evolved towards purely attention mechanism-based models [3]. Nowadays, state-of-the-art models on image retrieval tasks only exhibit end-to-end DNN-based architectures.
Despite the strides made by previous deep hashing-based methods using gradient-based end-to-end deep learning models with specifically designed loss functions to capture global and local information of images, this black-box training approach does not guarantee the embedding of desired information, including local or spatial details. Furthermore, because these end-to-end deep learning models are trained with predetermined criteria, they have fundamental limitations in conducting image retrieval flexibly, such as with additional conditions like precise positioning of each object or prioritizing specific objects during image retrieval.
To resolve the above limitations of previous methods, we propose an innovative image hashing method employing Hyperdimensional Computing (HDC) [14] to facilitate image retrieval with spatial structural conditions that can be easily manipulated as illustrated in Figure 1. HDC stands as an alternative paradigm inspired by essential brain functions, emphasizing high efficiency and symbolic learning capabilities. Recently, HDC has demonstrated the power of combining with neural network models [11]. Building on the success of neuro-symbolic AI with HDC and the observation that the human brain excels at manipulating high-dimensional representations, our approach harnesses HDC operations to embed spatial structural information into a high-dimensional vector in a neuro-symbolic manner, constituting a hashed representation of the image.
Our methodology capitalizes on pre-trained large vision models to extract feature representations for individual objects, subsequently combining them into a singular representation of high-dimensional vectors through spatial encoding – a process applying HDC operations with positional information. These representations are then hashed into lower-dimensional bipolar vectors to facilitate rapid image retrieval. During the retrieval process, our method replicates the spatial encoding procedure to retrieve images with similar spatial structures. Additionally, structural conditions can be controlled by incorporating HDC operations on a given query image, such as focusing on spatial information of a specific object.
In summary, our work represents a fundamentally novel contribution to the field, offering the following key advancements:
- •
We propose NeuroHash , a novel self-supervised neuro-symbolic image hashing framework designed to enable customizable spatial-aware image retrieval. Unlike previous methods that rely on fully gradient-based end-to-end deep learning models, our approach combines DNN-based models with HDC-based symbolic models to symbolically encode spatial information with local features, allowing for flexible manipulation of hash values and enhancing interpretability. This enables conditional image retrieval that can focus on specific objects or spatial regions within an image, offering a high degree of customization and control.
- •
We introduce a hyperdimensional spatial encoding technique, which is, to the best of our knowledge, the first HDC encoding method that preserves spatial similarity.
- •
We devise a context-aware HDC encoder that preserves characteristics from the original feature space using self-supervised training, advancing previous HDC encoders.
- •
We propose new loss terms for lower-dimensional bipolar hashing using multilinear hyperplanes to enhance hash function optimization.
- •
Experimental results on two benchmark datasets show that NeuroHash outperforms state-of-the-art hashing methods in terms of mAP@5K scores. These results demonstrate that our neuro-symbolic framework can achieve performance comparable to fully gradient-based end-to-end models while offering additional benefits in terms of flexibility and customization.
- •
We introduce a new metric, mAP@5Kr, to evaluate the effectiveness of spatial-aware and conditional image retrieval, demonstrating the framework’s capability to align retrieved images with spatial constraints accurately.
II Related Works
II-A Hash-based Approximate Nearest Search
Retrieving similar vectors efficiently from abundant vector data using linear search or traditional structures is impractical. To address this, studies explore converting high-dimensional vectors into fixed-size, low-dimensional representations, with Locality Sensitive Hashing (LSH)[8] being a notable unsupervised algorithm[30]. LSH constructs a hash table using multiple functions capturing local similarity, and variations like Multilinear Hyperplane Hashing [24] specifically preserve cosine similarity in the hash space.
II-B Deep Hashing Approach
In early image retrieval, methods such as supervised discrete hashing (SDH)[29] played a crucial role in reducing storage and improving retrieval speed. The integration of Convolutional Neural Networks (CNNs) brought advancements with models such as HashNet [2], building on architectures like AlexNet [17] to address discrete optimization challenges.
The evolution shifted towards deep learning, leveraging ResNet as a popular backbone network in approaches like CSQ [41] and DBDH [42]. As models progressed, attention turned to hybrid models like DAgH [4] and DAHP [18], using attention networks to enhance performance without increasing convolution layers. Scalability concerns led to exploration of a self-attention-based structure [3].
Recent developments expanded unsupervised deep hashing into applications like image copy detection [21] and image quality assessment [12]. Methods like DeepBit [19], DistillHash [39], and TBH [32] explored unsupervised learning with novel loss functions. Contrastive learning in computer vision paved the way for unsupervised hashing methods such as HAMAN [25] and MeCoQ [35], leveraging contrasting positive and negative samples for robust hash codes.
Certain unsupervised hashing methods focused on mining pairwise similarity. DistillHash [39] and SSDH [38] used data pair distillation and semantic structures, while FSCH [1] extended these approaches with fine-grained similarity structures based on global and local image representations. Our proposed approach takes a completely different direction by leveraging an HDC-enhanced neuro-symbolic approach while following the previous strategy that combines local and global features.
II-C Hyperdimensional Computing
Brain-inspired hyperdimensional computing (HDC) is based on the understanding that brains compute with patterns of neural activity that are not readily associated with numbers. Due to the huge size of the brain’s circuits, neural patterns can be modeled with hypervectors [14]. HDC builds upon a well-defined set of operations with random hypervectors, is extremely robust in the presence of failures, and offers a complete computational paradigm that is easily applied to multiple learning problems, such as speech recognition [13], graph learning [15, 26], and computer vision [10, 7].
Recent literature has witnessed a growing interest in hyperdimensional computing (HDC) as a learning model, praised for its simplicity and computational efficiency. However, conventional HDC frameworks encounter issues with randomly generated and static encoders, leading to an abundance of parameters and decreased accuracy. LeHD [6], an innovative approach, employs a principled learning approach to refine model accuracy, transforming the HDC framework into an equivalent binary neural network architecture. These advancements collectively aim to overcome issues with static encoders in HDC, offering a more effective and accurate learning framework.
III Methodology
III-A HDC Basics
The core of HDC is called a hyperdimensional vector, denoted , which represents a vector in with a high dimensionality of . Hyperdimensional vectors are compared using a similarity function . By using this similarity measure, HDC becomes a versatile tool for cognitive tasks, including memory, classification, clustering, etc. HDC frameworks designed to support these tasks are based on three core operations that mirror brain functionalities: bundling, binding, and permutation. Here are the details of each operation:
- 1.
Bundling: This operation, represented by , is commonly executed as element-wise addition. If , then both and exhibit similarity to . In terms of cognitive interpretation, this operation can be understood as a form of memorization.
- 2.
Binding: This operation, denoted by , is usually implemented as an element-wise multiplication. If , then is dissimilar to both and . Binding has a crucial property of similarity preservation, where for some hypervector , . From a cognitive point of view, this operation can be understood as an association. Binding can be used to associate different pieces of information, such as coordinates and image feature vectors, in hyperdimensional space.
- 3.
Permutation: This operator, represented by , is commonly executed as a rotation of vector elements. In general, . Permutation is frequently employed to encode the order within sequences.
Leveraging the three fundamental HDC operations provides a foundation for a hyperdimensional learning framework applicable to various tasks. In the context of classification, each step of the framework can be outlined as follows.
- 1.
Encoding: The initial step within the HDC framework involves mapping the input data into a high-dimensional space through the introduction of an encoding function , commonly known as encoding. Consider an input vector with features, denoted as , representing features extracted from an image. The commonly used encoding function is defined as , where is an matrix, and each element in is sampled from an i.i.d Gaussian distribution with parameters (). Additionally, is sampled from an i.i.d uniform distribution over the interval . The function preserves a notion of similarity in the input space. Consequently, for any given inputs , their corresponding hypervectors, and , exhibit similarity iif is similar to . Such initialized encoders with parameters and can be further optimized by making and learnable parameters using a gradient descent approach.
- 2.
Symbolic Training: Consider a dataset where each data point is associated with a label from a set of classes. In traditional hyperdimensional classifier training, the process involves generating class hypervectors through bundling: . For each data point to retrain , each class hypervector is updated as follows:
where , , and is learning rate.
- 3.
Symbolic Inference: Once the class hypervectors undergo updates through the initial training phases, the classification of a given query becomes a straightforward process. A class is predicted when is satisfied for all .
III-B Proposed Framework
III-B1 Overall Pipeline
The overall pipeline of our proposed framework is presented in Figure 2. First, given an image , we extract global features, which is embedding of the image, through a pre-trained image encoder model by giving the entire given image to the model (). Also, in order to consider local information, it extracts bounding boxes indicating objects that are presented in the image by conducting an object detection task over the image using a pre-trained object detection model (
III-B2 Global and Local Visual Features Extraction
In an image retrieval task, it is crucial to well-represent each image in a compact representation. Although pre-trained large image embedding models introduced so far present powerful performance in extracting visual features, simply embedding entire images can lead to insufficient interpretation of local information considering the complexity of image data. To allow solid local visual information consideration, we propose to employ a pre-trained object detection model in order to extract objects that are presented in a given image. Therefore, our proposed framework uses two pre-trained large image models: 1) object detection model
III-B3 Context-aware HDC Encoding
Inspired by the previous work LeHD [6], we designed an HDC encoder
| (1) |
| (2) |
III-B4 Hyperdimensional Spatial-aware Encoding
Given global feature hypervector
Additionally, we can introduce a new hyperparameter length scale
Now, to have the final hyperdimensional representation, positional hypervectors are combined with the visual feature hypervectors that are retrieved from the global and local visual features extraction process. Each local visual feature vector
Furthermore, we can utilize Symbolic Training shown in subsection III-A where we merge separate symbolic representations into a single hypervector and optimize by giving weights to each symbolic hypervector to have a user desire hyperdimensional representations:
III-B5 Multilinear Hyperplane Hashing Optimization
We explored that by utilizing HDC operations in hyperspace, hyperdimensional representation
To have a well-performing hash function
To optimize randomly sampled hyperplanes from a normal distribution, we generalized our hashing function as
| (3) |
The loss function that is shown in Equation 3 consists of 5 loss terms: mean square error (MSE) loss
Methods
References
CIFAR10
MS COCO
16 bits
32 bits
64 bits
16 bits
32 bits
64 bits
AGH [23]
ICML11
0.333
0.357
0.358
0.596
0.625
0.631
ITQ [9]
TPAMI12
0.305
0.325
0.349
0.598
0.624
0.648
DGH [22]
NeurIPS14
0.335
0.353
0.361
0.613
0.631
0.638
SGH [5]
ICML17
0.435
0.437
0.433
0.594
0.610
0.618
BGAN [33]
AAAI18
0.525
0.531
0.562
0.645
0.682
0.707
GreedyHash [34]
NeurIPS18
0.448
0.473
0.501
0.582
0.668
0.710
DVB [31]
IJCV19
0.403
0.422
0.446
0.570
0.629
0.623
TBH [32]
CVPR20
0.497
0.524
0.529
0.706
0.735
0.722
CIB [28]
IJCAI21
0.547
0.583
0.602
0.737
0.760
0.775
HAMAN [25]
IJCAI22
-
-
-
0.722
0.775
0.787
NSH [40]
IJCAI22
0.706
0.733
0.756
0.746
0.774
0.783
FSCH [1]
TCSVT23
0.876
0.912
0.926
0.760
0.787
0.799
naïve (DINOv2 + LSH)
0.316
0.450
0.599
0.479
0.557
0.658
NeuroHash (
Without
mAP@5K
Loss term for numerical correspondence.
First of all, the MSE loss term
| (4) |
To match the hamming distance value with the similarity value, we used reversed hamming distance:
Loss terms for limited representation.
Due to the low precision bits representation, distance is also extremely discrete which makes indistinguishable distances between many images. To tackle this issue, we set an assumption that in most cases, boundary distance is not placed among distances that are located on either side of the edges – either distance is very close or very far. Based on this assumption, we applied another loss term we named w-shape loss presented in Equation 5. This loss function gives more penalty for the distances that are more closely located in the center. In the same context of low precision and low dimensionality, it also can cause limited unique representations. To avoid such representation collapsing, we also introduced uniform loss as shown in Equation 6. Note that
| (5) |
| (6) |
Loss term for learning binary representations.
Next, since we are using
| (7) |
Loss term for reversed relative order.
For the last loss term, we consider the relative orders between hyperdimensional representation pairs. It targets to preserve the order of ranking that each
| (8) |
IV Experiments
IV-1 Experiment settings
Scale Factor (
- 1.
Implementation Details For the object detection model
we used Detectron 2 [36] and for the image embedding modelf o b j ( . ) f_{obj}(.) we used DINOv2 ViT-g/14 model [27]. Since the ViT-g/14 model uses a patch size of 14, it is necessary to transform the image size into a multiplier of 14 for both width and height. It is implemented as transforming a given imageϕ → v i s ( . ) \vec{\phi}_{vis}(.) of sizeI k I_{k} to( w , h ) (w,h) . For the hypervectors, we used the dimensionality of( ( ⌊ w / 14 ⌋ + 1 ) × 14 , ( ⌊ h / 14 ⌋ + 1 ) × 14 ) ((\lfloor{w/14}\rfloor+1)\times 14,(\lfloor{h/14}\rfloor+1)\times 14) , which is commonly selected in the HDC domain.D = 10,000 D=10,000 - 2.
Evaluation Metrics In order to thoroughly evaluate our proposed method and compare it to conventional baselines, we used mAP (mean Average Precision), a widely accepted metric for evaluating retrieval performance. This metric calculates the average precision (AP) for a given query and a ranked list of returned results, where mAP is determined by averaging the AP values across all queries. In our evaluation, we follow the latest convention and use mAP@5000 for CIFAR-10 and MS COCO. Higher mAP values indicate better overall performance.
In addition to conventional evaluation metrics, we introduce a novel metric called mAP@K
to measure the effectiveness of our proposed spatial-aware conditional image retrieval. This metric represents a spatial-aware version of mAP and evaluates whether the coordinates of objects in the query image align with those in the retrieved image. This alignment is determined by calculating the Euclidean distance between the ground truth object coordinates and those of the retrieved image. The parameterr {r} defines the metric’s spatial sensitivity by determining correct retrieval for two objects’r r andi i having the same class usingj j where( x i w i − x j w j ) 2 + ( y i h i − y j h j ) 2 ≤ r 2 (\frac{x_{i}}{w_{i}}-\frac{x_{j}}{w_{j}})^{2}+(\frac{y_{i}}{h_{i}}-\frac{y_{j}}{h_{j}})^{2}\leq r^{2} andx i , y i x_{i},y_{i} represent each object’s coordinates in their image andx j , x j x_{j},x_{j} andw i , h i w_{i},h_{i} indicate each image’s dimensionality. Consequently, a higher value ofw j , h j w_{j},h_{j} results in a more lenient evaluation of whether the retrieved object contains similar objects at the same location. As the coordinate information is crucial for our proposed framework, we performed evaluations exclusively on the MS COCO dataset, using mAP@Kr r , mAP@Kr = 0.1 {r=0.1} , mAP@Kr = 0.2 {r=0.2} and mAP@Kr = 0.3 {r=0.3} .r = 0.4 {r=0.4}
IV-2 Datasets
- 1.
MS-COCO [20] has 82,783 training samples and 40,504 validation samples, with each image annotated with one or more labels from a pool of 91 categories. In this study, we follow the previous research [1], a subset of 122,218 images from 80 categories is used. Within this subset, a random sample of 5,000 images is referred to as the query dataset, while the remaining images form the retrieval set. In particular, MS COCO stands out from other datasets due to the inclusion of ground truth bounding box information, providing a unique opportunity to assess the extent to which our proposed method captures local information using our proposed metric mAP@K
.r {r} - 2.
CIFAR-10 [16] involves 60,000 images distributed across 10 categories, with each class containing 6,000 images. Following the earlier study [1], we randomly chose 100 images from each class to form the query dataset, amounting to a total of 1000 images. Subsequently, we utilized the remaining images for retrieval purposes.
IV-3 Evaluation on Weak-spatial-aware Image Retrieval
First, we evaluated our NeuroHash on Weak-spatial-aware image retrieval case with other hashing methods including current state-of-the-art models. On Weak-spatial-aware image retrieval, we focus on conventional image retrieval metric mAP@K which evaluates without spatial alignment of each object shown in the images. In this test, we gave high-scale factors
IV-4 Evaluation on Strong-spatial-aware Image Retrieval
To ensure the efficacy of our proposed NeuroHash on spatial-aware conditional image retrieval task, we conducted image retrieval evaluation on Strong-spatial-aware image retrieval case which aims to retrieve images with similar object positioning. In this evaluation, we use our proposed mAP@5Kr metric with
IV-5 Evaluation on Conditional Image Retrieval
In this conditional image retrieval section, we visually demonstrate spatial-aware image retrieval and conditional retrieval of our NeuroHash shown in Figure 3. On Figure 3.(a) shows the effect of controlling
IV-6 Ablation Study on Multilinear Hyperplane Hashing Opt.
Table IIshows an ablation study on our model. The results indicate that all loss metrics are necessary for effective hash value generation on the CIFAR10 dataset using the mAP@5K metric. As shown, the full model achieved the highest score.
V Conclusions
In this paper, we propose NeuroHash a completely novel approach to hashing images in a neuro-symbolic way that enables spatial-aware hashing and conditional image retrieval. Experiments on well-known datasets for image retrieval performance benchmarking validate the efficacy of our work. In future work, we aim to evolve a more versatile approach capable of embedding various types of information, including temporal information. This future work seeks to broaden the scope of our neuro-symbolic framework, fostering its application in diverse domains beyond spatial-aware image retrieval.
Acknowledgements
This work was supported in part by the DARPA Young Faculty Award, the National Science Foundation (NSF) under Grants #2127780, #2319198, #2321840, #2312517, and #2235472, the Semiconductor Research Corporation (SRC), the Office of Naval Research through the Young Investigator Program Award, and Grants #N00014-21-1-2225 and #N00014-22-1-2067. Additionally, support was provided by the Air Force Office of Scientific Research under Award #FA9550-22-1-0253, along with generous gifts from Xilinx and Cisco.
References
- [1] H. Cao, L. Huang, J. Nie, and Z. Wei, “Unsupervised deep hashing with fine-grained similarity-preserving contrastive learning for image retrieval,” IEEE Transactions on Circuits and Systems for Video Technology, 2023.
- [2] Z. Cao, M. Long, J. Wang, and P. S. Yu, “Hashnet: Deep learning to hash by continuation,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 5608–5617.
- [3] C.-F. R. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 357–366.
- [4] Y. Chen, Z. Lai, Y. Ding, K. Lin, and W. K. Wong, “Deep supervised hashing with anchor graph,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9796–9804.
- [5] B. Dai, R. Guo, S. Kumar, N. He, and L. Song, “Stochastic generative hashing,” in International Conference on Machine Learning. PMLR, 2017, pp. 913–922.
- [6] S. Duan, Y. Liu, S. Ren, and X. Xu, “Lehdc: Learning-based hyperdimensional computing classifier,” in Proceedings of the 59th ACM/IEEE Design Automation Conference, 2022, pp. 1111–1116.
- [7] A. Dutta, S. Gupta, B. Khaleghi, R. Chandrasekaran, W. Xu, and T. Rosing, “Hdnn-pim: Efficient in memory design of hyperdimensional computing with feature extraction,” in Proceedings of the Great Lakes Symposium on VLSI 2022, 2022, pp. 281–286.
- [8] A. Gionis, P. Indyk, R. Motwani et al., “Similarity search in high dimensions via hashing,” in Vldb, vol. 99, no. 6, 1999, pp. 518–529.
- [9] Y. Gong, S. Lazebnik, A. Gordo, and F. Perronnin, “Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval,” IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 12, pp. 2916–2929, 2012.
- [10] M. Hersche, G. Karunaratne, G. Cherubini, L. Benini, A. Sebastian, and A. Rahimi, “Constrained few-shot class-incremental learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9057–9067.
- [11] M. Hersche, M. Zeqiri, L. Benini, A. Sebastian, and A. Rahimi, “A neuro-vector-symbolic architecture for solving raven’s progressive matrices,” Nature Machine Intelligence, vol. 5, no. 4, pp. 363–375, 2023.
- [12] Z. Huang and S. Liu, “Perceptual hashing with visual content understanding for reduced-reference screen content image quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 7, pp. 2808–2823, 2020.
- [13] M. Imani, D. Kong, A. Rahimi, and T. Rosing, “Voicehd: Hyperdimensional computing for efficient speech recognition,” in 2017 IEEE international conference on rebooting computing (ICRC). IEEE, 2017, pp. 1–8.
- [14] P. Kanerva, “Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional vectors,” Cognitive Computation, 2009.
- [15] J. Kang, M. Zhou, A. Bhansali, W. Xu, A. Thomas, and T. Rosing, “Relhd: A graph-based learning on fefet with hyperdimensional computing,” in 2022 IEEE 40th International Conference on Computer Design (ICCD). IEEE, 2022, pp. 553–560.
- [16] A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009.
- [17] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems, vol. 25, 2012.
- [18] X. Li, J. Yu, Y. Wang, J.-Y. Chen, P.-X. Chang, and Z. Li, “Dahp: Deep attention-guided hashing with pairwise labels,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 3, pp. 933–946, 2021.
- [19] K. Lin, J. Lu, C.-S. Chen, J. Zhou, and M.-T. Sun, “Unsupervised deep learning of compact binary descriptors,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 6, pp. 1501–1514, 2018.
- [20] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. Springer, 2014, pp. 740–755.
- [21] S. Liu and Z. Huang, “Efficient image hashing with geometric invariant vector distance for copy detection,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 15, no. 4, pp. 1–22, 2019.
- [22] W. Liu, C. Mu, S. Kumar, and S.-F. Chang, “Discrete graph hashing,” Advances in neural information processing systems, vol. 27, 2014.
- [23] W. Liu, J. Wang, S. Kumar, and S.-F. Chang, “Hashing with graphs,” 2011.
- [24] X. Liu, X. Fan, C. Deng, Z. Li, H. Su, and D. Tao, “Multilinear hyperplane hashing,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5119–5127.
- [25] Z. Ma, W. Ju, X. Luo, C. Chen, X.-S. Hua, and G. Lu, “Improved deep unsupervised hashing via prototypical learning,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 659–667.
- [26] I. Nunes, M. Heddes, T. Givargis, A. Nicolau, and A. Veidenbaum, “Graphhd: Efficient graph classification using hyperdimensional computing,” in 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2022, pp. 1485–1490.
- [27] M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P.-Y. Huang, H. Xu, V. Sharma, S.-W. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski, “Dinov2: Learning robust visual features without supervision,” 2023.
- [28] Z. Qiu, Q. Su, Z. Ou, J. Yu, and C. Chen, “Unsupervised hashing with contrastive information bottleneck,” arXiv preprint arXiv:2105.06138, 2021.
- [29] F. Shen, C. Shen, W. Liu, and H. Tao Shen, “Supervised discrete hashing,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 37–45.
- [30] F. Shen, Y. Xu, L. Liu, Y. Yang, Z. Huang, and H. T. Shen, “Unsupervised deep hashing with similarity-adaptive and discrete optimization,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 12, pp. 3034–3044, 2018.
- [31] Y. Shen, L. Liu, and L. Shao, “Unsupervised binary representation learning with deep variational networks,” International Journal of Computer Vision, vol. 127, no. 11-12, pp. 1614–1628, 2019.
- [32] Y. Shen, J. Qin, J. Chen, M. Yu, L. Liu, F. Zhu, F. Shen, and L. Shao, “Auto-encoding twin-bottleneck hashing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2818–2827.
- [33] J. Song, T. He, L. Gao, X. Xu, A. Hanjalic, and H. T. Shen, “Binary generative adversarial networks for image retrieval,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018.
- [34] S. Su, C. Zhang, K. Han, and Y. Tian, “Greedy hash: Towards fast optimization for accurate hash coding in cnn,” Advances in neural information processing systems, vol. 31, 2018.
- [35] J. Wang, Z. Zeng, B. Chen, T. Dai, and S.-T. Xia, “Contrastive quantization with code memory for unsupervised image retrieval,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 2468–2476.
- [36] Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2, 2019.
- [37] R. Xia, Y. Pan, H. Lai, C. Liu, and S. Yan, “Supervised hashing for image retrieval via image representation learning,” in Proceedings of the AAAI conference on artificial intelligence, vol. 28, no. 1, 2014.
- [38] E. Yang, C. Deng, T. Liu, W. Liu, and D. Tao, “Semantic structure-based unsupervised deep hashing,” in Proceedings of the 27th international joint conference on artificial intelligence, 2018, pp. 1064–1070.
- [39] E. Yang, T. Liu, C. Deng, W. Liu, and D. Tao, “Distillhash: Unsupervised deep hashing by distilling data pairs,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 2946–2955.
- [40] J. Yu, Y. Shen, M. Wang, H. Zhang, and P. H. Torr, “Learning to hash naturally sorts,” arXiv preprint arXiv:2201.13322, 2022.
- [41] L. Yuan, T. Wang, X. Zhang, F. E. Tay, Z. Jie, W. Liu, and J. Feng, “Central similarity quantization for efficient image and video retrieval,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 3083–3092.
- [42] X. Zheng, Y. Zhang, and X. Lu, “Deep balanced discrete hashing for image retrieval,” Neurocomputing, vol. 403, pp. 224–236, 2020.