<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Shahid Rajaee Teacher Training University</PublisherName>
				<JournalTitle>Journal of Electrical and Computer Engineering Innovations (JECEI)</JournalTitle>
				<Issn>2322-3952</Issn>
				<Volume>12</Volume>
				<Issue>1</Issue>
				<PubDate PubStatus="epublish">
					<Year>2024</Year>
					<Month>01</Month>
					<Day>01</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Text Detection and Recognition for Robot Localization</ArticleTitle>
<VernacularTitle></VernacularTitle>
			<FirstPage>163</FirstPage>
			<LastPage>174</LastPage>
			<ELocationID EIdType="pii">1986</ELocationID>
			
<ELocationID EIdType="doi">10.22061/jecei.2023.9857.658</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>Z.</FirstName>
					<LastName>Raisi</LastName>
<Affiliation>University of Waterloo, Waterloo, Canada and
Chabahar Maritime University, Chabahar, Iran.</Affiliation>

</Author>
<Author>
					<FirstName>J.</FirstName>
					<LastName>Zelek</LastName>
<Affiliation>Systems Design Engineering Department, University of Waterloo, Canada.</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2023</Year>
					<Month>06</Month>
					<Day>26</Day>
				</PubDate>
			</History>
		<Abstract>&lt;strong&gt;Background and Objectives:&lt;/strong&gt; Signage is everywhere, and a robot should be able to take advantage of signs to help it localize (including Visual Place Recognition (VPR)) and map. Robust text detection &amp; recognition in the wild is challenging due to pose, irregular text instances, illumination variations, viewpoint changes, and occlusion factors.&lt;br /&gt;&lt;strong&gt;Methods: &lt;/strong&gt;This paper proposes an end-to-end scene text spotting model that simultaneously outputs the text string and bounding boxes. The proposed model leverages a pre-trained Vision Transformer based (ViT) architecture combined with a multi-task transformer-based text detector more suitable for the VPR task. Our central contribution is introducing an end-to-end scene text spotting framework to adequately capture the irregular and occluded text regions in different challenging places. We first equip the ViT backbone using a masked autoencoder (MAE) to capture partially occluded characters to address the occlusion problem. Then, we use a multi-task prediction head for the proposed model to handle arbitrary shapes of text instances with polygon bounding boxes.&lt;br /&gt;&lt;strong&gt;Results:&lt;/strong&gt; The evaluation of the proposed architecture&#039;s performance for VPR involved conducting several experiments on the challenging Self-Collected Text Place (SCTP) benchmark dataset. The well-known evaluation metric, Precision-Recall, was employed to measure the performance of the proposed pipeline. The final model achieved the following performances, Recall = 0.93 and Precision = 0.8, upon testing on this benchmark.&lt;br /&gt;&lt;strong&gt;Conclusion:&lt;/strong&gt; The initial experimental results show that the proposed model outperforms the state-of-the-art (SOTA) methods in comparison to the SCTP dataset, which confirms the robustness of the proposed end-to-end scene text detection and recognition model.</Abstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">Text detection</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Text Recognition</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Robotics Localization</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Deep Learning</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Visual Place Recognition</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://jecei.sru.ac.ir/article_1986_d84532fa293f8b6440ac046644e32bb0.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
