题目

All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACGAATTCCG". When studying DNA, it is sometimes useful to identify repeated sequences within the DNA.

Write a function to find all the 10-letter-long sequences (substrings) that occur more than once in a DNA molecule.

原题链接：https://oj.leetcode.com/problems/repeated-dna-sequences/

straight-forward method（TLE）

算法分析

直接字符串匹配；设计next数组，存字符串中每个字母在其中后续出现的位置；遍历时以next数组为起始。

简化考虑长度为4的字符串

case1:

src A C G T A C G T

next [4] [5] [6] [7] [-1] [-1] [-1] [-1]

那么匹配ACGT字符串的过程，匹配next[0]之后的3位字符即可

case2：

src A C G T A A C G T

next [4] [5] [6] [7] [5] [-1] [-1] [-1] [-1]

多个A字符后继，那么需要匹配所有后继，匹配next[0]不符合之后，还要匹配next[next[0]]

case3：

src A A A A A A

next [1] [2] [3] [4] [5] [-1]

重复的情况，在next[0]匹配成功时，可以把next[next[0]]置为-1，即以next[0]开始的长度为4的字符串已经成功匹配过了，无需再次匹配了；当然这么做只能减少重复的情况，并不能消除重复，因此仍需要使用一个set存储匹配成功的结果，方便去重

时间复杂度

构造next数组的复杂度O(n^2)，遍历的复杂度O(n^2)；总时间复杂度O(n^2)

代码实现

 #include <string>

 #include <vector>

 #include <set>

 class Solution {

 public:

     std::vector<std::string> findRepeatedDnaSequences(std::string s);

     ~Solution();

 private:

     std::size_t* next;

 };

 std::vector<std::string> Solution::findRepeatedDnaSequences(std::string s) {

     std::vector<std::string> rel;

     if (s.length() <= ) {

         return rel;

     }

     next = new std::size_t[s.length()];

     // cal next array

     for (int pos = ; pos < s.length(); ++pos) {

         next[pos] = s.find_first_of(s[pos], pos + );

     }

     std::set<std::string> tmpRel;

     for (int pos = ; pos < s.length(); ++pos) {

         std::size_t nextPos = next[pos];

         while (nextPos != std::string::npos) {

             int ic = pos;

             int in = nextPos;

             int count = ;

             while (in != s.length() && count <  && s[++ic] == s[++in]) {

                 ++count;

             }

             if (count == ) {

                 tmpRel.insert(s.substr(pos, ));

                 next[nextPos] = std::string::npos;

             }

             nextPos = next[nextPos];

         }

     }

     for (auto itr = tmpRel.begin(); itr != tmpRel.end(); ++itr) {

         rel.push_back(*itr);

     }

     return rel;

 }

 Solution::~Solution() {

     delete [] next;

 }

hash table plus bit manipulation method

（view the Show Tags and Runtime 10ms !）

算法分析

首先考虑将ACGT进行二进制编码

A -> 00

C -> 01

G -> 10

T -> 11

在编码的情况下，每10位字符串的组合即为一个数字，且10位的字符串有20位；一般来说int有4个字节，32位，即可以用于对应一个10位的字符串。例如

ACGTACGTAC -> 00011011000110110001

AAAAAAAAAA -> 00000000000000000000

20位的二进制数，至多有2^20种组合，因此hash table的大小为2^20，即1024 * 1024，将hash table设计为bool hashTable[1024 * 1024];

遍历字符串的设计

每次向右移动1位字符，相当于字符串对应的int值左移2位，再将其最低2位置为新的字符的编码值，最后将高2位置0。例如

src CAAAAAAAAAC

subStr CAAAAAAAAA

int 0100000000

subStr AAAAAAAAAC

int 0000000001

时间复杂度

字符串遍历O(n)，hash tableO(1)；总时间复杂度O(n)

代码实现

 #include <string>

 #include <vector>

 #include <unordered_set>

 #include <cstring>

 bool hashMap[*];

 class Solution {

 public:

     std::vector<std::string> findRepeatedDnaSequences(std::string s);

 };

 std::vector<std::string> Solution::findRepeatedDnaSequences(std::string s) {

     std::vector<std::string> rel;

     if (s.length() <= ) {

         return rel;

     }

     // map char to code

     unsigned char convert[];

     convert[] = ; // 'A' - 'A'  00

     convert[] = ; // 'C' - 'A'  01

     convert[] = ; // 'G' - 'A'  10

     convert[] = ; // 'T' - 'A' 11

     // initial process

     // as ten length string

     memset(hashMap, false, sizeof(hashMap));

     int hashValue = ;

     for (int pos = ; pos < ; ++pos) {

         hashValue <<= ;

         hashValue |= convert[s[pos] - 'A'];

     }

     hashMap[hashValue] = true;

     std::unordered_set<int> strHashValue;

     //

     for (int pos = ; pos < s.length(); ++pos) {

         hashValue <<= ;

         hashValue |= convert[s[pos] - 'A'];

         hashValue &= ~(0x300000);

         if (hashMap[hashValue]) {

             if (strHashValue.find(hashValue) == strHashValue.end()) {

                 rel.push_back(s.substr(pos - , ));

                 strHashValue.insert(hashValue);

             }

         } else {

             hashMap[hashValue] = true;

         }

     }

     return rel;

 }

Leetcode：Repeated DNA Sequences详细题解的更多相关文章

[LeetCode] Repeated DNA Sequences 求重复的DNA序列
All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACG ...
[Leetcode] Repeated DNA Sequences
All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACG ...
LeetCode() Repeated DNA Sequences 看的非常的过瘾！
All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACG ...
[LeetCode] Repeated DNA Sequences hash map
All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACG ...
lc面试准备:Repeated DNA Sequences
1 题目 All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: &quo ...
LeetCode 187. 重复的DNA序列(Repeated DNA Sequences)
187. 重复的DNA序列 187. Repeated DNA Sequences 题目描述 All DNA is composed of a series of nucleotides abbrev ...
【LeetCode】Repeated DNA Sequences 解题报告
[题目] All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: &quo ...
[LeetCode] 187. Repeated DNA Sequences 求重复的DNA序列
All DNA is composed of a series of nucleotides abbreviated as A, C, G, and T, for example: "ACG ...
【LeetCode】187. Repeated DNA Sequences 解题报告（Python）
作者: 负雪明烛 id: fuxuemingzhu 个人博客: http://fuxuemingzhu.cn/ 题目地址: https://leetcode.com/problems/repeated ...

随机推荐

Linux 常用命令使用方法大搜刮(转)
1.# 表示权限用户(如:root),$ 表示普通用户开机提示:Login:输入用户名 password:输入口令用户是系统注册用户成功登陆后,可以进入相应的用户环境. 退出当前shel ...
Qt 学习之路：Canvas
在 QML 刚刚被引入到 Qt 4 的那段时间,人们往往在讨论 Qt Quick 是不是需要一个椭圆组件.由此,人们又联想到,是不是还需要其它的形状?这种没玩没了的联想导致了一个最直接的结果:除了圆角 ...
js原型继承
原型链: Object(构造函数) object(类型(对象)) var o = {}; alert(typeof o); //结果是object alert(typeof Object); //结果 ...
CUDA与VS2013安装
@import url(http://i.cnblogs.com/Load.ashx?type=style&file=SyntaxHighlighter.css);@import url(/c ...
spring事务管理学习
spring事务管理学习 spring的事务管理和mysql自己的事务之间的区别参考很好介绍事务异常回滚的文章 MyBatis+Spring 事务管理 spring中的事务回滚例子这篇文章讲解了@ ...
ASP.NET和支付宝合作开发第三方接口的注意事项
最近公司和支付宝合作开发第三方接口的项目,这里把过程中需要注意的地方说明一下: 前提:一般来说单个银行不接收个人或私企开通支付接口.因此,和第三方支付公司合作,签订合约开放接口就是通行的做法. 流程: ...
头一回发博客,来分享个有关C++类型萃取的编写技巧
废话不多说,上来贴代码最实在,哈哈! 以下代码量有点多,不过这都是在下一手一手敲出来的,小巧好用,把以下代码复制出来,放到相应的hpp文件即可,VS,GCC下均能编译通过 #include<io ...
代码bug
1.webstorm ide未配置basePath本地会加入根路径 2.点击一次就销毁可以给标签设置一个值data-val="0" 某个函数只执行一次的方法,或者也可以考虑绑用on ...
Overloads和Overrides在元属性继承上的特性
元属性继承可以使用IsDefined函数进行判断,先写出结论如果使用Overrides,则元属性可以继承,除非在使用IsDefined时明确不进行继承判断,如 pFunction.IsDefined ...
Html5新增加的属性
用2中方法给单复选框增加新的特性,使直接点击文字就可以被选中 1.将选项放入label标签内添加for属性,并在input标签内添加id,两者值相同. 2.将input标签放到label标签内,注意l ...

Leetcode：Repeated DNA Sequences详细题解

题目

straight-forward method（TLE）

算法分析

case1:

case2：

case3：

时间复杂度

代码实现

hash table plus bit manipulation method

（view the Show Tags and Runtime 10ms !）

算法分析

遍历字符串的设计

时间复杂度

代码实现

Leetcode：Repeated DNA Sequences详细题解的更多相关文章

随机推荐

热门专题